publications

publications by categories in reversed chronological order. generated by jekyll-scholar.

2026

  1. SIGCOMM26
    Connex: Endpoint Mobility Primitives for Dynamic LLM Serving
    Yanying Lin, Vincent Liu, Tao Luo, Chengzhong Xu, and Kejiang Ye
    In Proceedings of the ACM SIGCOMM Conference, 2026
  2. ISCA26
    DynoPipe: Heterogeneous Edge-Cloud LLM Serving with Dynamically Orchestrated Pipeline Boundaries
    Yanying Lin, Baicheng Chen, Xinyu Zhang, Chengzhong Xu, and Kejiang Ye
    In Proceedings of the 53rd International Symposium on Computer Architecture, 2026
  3. TPDS26
    Workload-Adapted Resource Allocation for LLM Distributed Serving in Serverless Clusters
    Yanying Lin, Shijie Peng, Yanbo Li, Shutian Luo, Haiying Shen, Kejiang Ye, and Chengzhong Xu
    IEEE Transactions on Parallel and Distributed Systems, 2026
  4. EuroSys26
    FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
    Yanying Lin, Shijie Peng, Chengzhi Lu, Chengzhong Xu, and Kejiang Ye
    In Proceedings of the 21st European Conference on Computer Systems, 2026
  5. ICPP26
    ReliefServe: Relieving GPU Pressure in Multi-Model Serving via Selective CPU Escape
    Shijie Peng, Yanying Lin*, Chengzhi Lu, Shuaipeng Wu, Chengzhong Xu, and Kejiang Ye
    In Proceedings of the 55th IEEE International Conference on Parallel Processing, 2026

2025

  1. SoCC25
    Understanding Diffusion Model Serving in Production: A Top-Down Analysis of Workload, Scheduling, and Resource Efficiency
    Yanying Lin, Shuaipeng Wu, Shutian Luo, Hong Xu, Haiying Shen, Chong Ma, Min Shen, Le Chen, Chengzhong Xu, Lin Qu, and Kejiang Ye
    In Proceedings of the 2025 ACM Symposium on Cloud Computing, 2025
  2. Cluster25
    ROCK: Serving Multimodal Model in Cloud with Heterogeneous-Aware Resource Orchestration for Thousands of LoRA Adapters
    Shuaipeng Wu, Yanying Lin*, Shijie Peng, Yanbo Li, Wenyan Chen, Chengzhong Xu, and Kejiang Ye
    In Proceedings of the 2025 IEEE International Conference on Cluster Computing, 2025
  3. IEEE TSC25
    Serving LLM in Distributed GPU Cluster with Fine-Grain Pipeline Constraints
    Yanying Lin, Shijie Peng, Shuaipeng Wu, Yanbo Li, Chengzhi Lu, Kejiang Ye, and Chengzhong Xu
    IEEE Transactions on Services Computing, 2025

2024

  1. ICDCS24
    Quart: Latency-Aware FaaS System for Pipelining Large Model Inference
    Yanying Lin, Yanbo Li, Shijie Peng, Yingfei Tang, Shutian Luo, Haiying Shen, Chengzhong Xu, and Kejiang Ye
    In Proceedings of the 44th IEEE International Conference on Distributed Computing Systems, 2024
  2. ICWS24
    Plank: Optimizing LLM Inference Performance in Pipeline Parallelism with Fine-Grained SLO Constraint
    Yanying Lin, Shijie Peng, Shuaipeng Wu, Yanbo Li, Chengzhi Lu, Chengzhong Xu, and Kejiang Ye
    In Proceedings of the 31st IEEE International Conference on Web Services, 2024
  3. IEEE TCC24
    Understanding Serverless Inference in Mobile-Edge Networks: a Benchmark Approach
    Junhong Chen, Yanying Lin*, Shijie Peng, Shuaipeng Wu, Kenneth Kent, Hao Dai, Kejiang Ye, and Yang Wang
    IEEE Transactions on Cloud Computing, 2024

2022

  1. IEEE TSC22
    Serverless Computing: State-of-the-Art, Challenges and Opportunities
    Yongkang Li, Yanying Lin, Kejiang Ye, Chengzhong Xu, and Yang Wang
    IEEE Transactions on Services Computing, 2022

2021

  1. IEEE/ACM ToN21
    A Novel End-to-end Deep Learning Framework for Encrypted Traffic Identification
    Peng Lin, Kejiang Ye, Yishen Hu, Yanying Lin, and Chengzhong Xu
    IEEE/ACM Transactions on Networking, 2021