Yanying Lin(林彦颖)

Postdoc, Harvard

yanyinglin.png

Postdoc@Harvard

Email: yylin1039@gmail.com

Yanying Lin is currently a postdoctoral researcher in Prof. Minlan Yu’s group at Harvard University. He received his Ph.D. in Computer Science from the University of Chinese Academy of Sciences (UCAS), where he worked with Prof. Kejiang Ye and Prof. Cheng-Zhong Xu. He previously visited the UPenn and UCSD. His research focuses on Heterogeneous-Native AI Infrastructure.

Research & Collaborations

Heterogeneous-Native AI Infrastructure addresses mixed accelerators, uneven networks, and dynamic AI workloads through communication, runtime, and HW/SW co-design: how tokens and activations cross network and device boundaries, how serving runtimes refactor pipelines as load and topology change, and how software is designed with the hardware that actually runs inference and agents.

He has collaborated with distinguished scholars in the systems community, including Prof. Rong Chen (SJTU), Prof. Xinyu Zhang (UCSD), Prof. Vincent Liu (UPenn), and Prof. Minlan Yu (Harvard).

Earlier visits took this into elastic agent inference at UPenn and FPGA-GPU acceleration for LLM serving at UCSD.

selected publications

  1. SIGCOMM26
    Connex: Endpoint Mobility Primitives for Dynamic LLM Serving
    Yanying Lin, Vincent Liu, Tao Luo, Chengzhong Xu, and Kejiang Ye
    In Proceedings of the ACM SIGCOMM Conference, 2026
  2. ISCA26
    DynoPipe: Heterogeneous Edge-Cloud LLM Serving with Dynamically Orchestrated Pipeline Boundaries
    Yanying Lin, Baicheng Chen, Xinyu Zhang, Chengzhong Xu, and Kejiang Ye
    In Proceedings of the 53rd International Symposium on Computer Architecture, 2026
  3. EuroSys26
    FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
    Yanying Lin, Shijie Peng, Chengzhi Lu, Chengzhong Xu, and Kejiang Ye
    In Proceedings of the 21st European Conference on Computer Systems, 2026
  4. SoCC25
    Understanding Diffusion Model Serving in Production: A Top-Down Analysis of Workload, Scheduling, and Resource Efficiency
    Yanying Lin, Shuaipeng Wu, Shutian Luo, Hong Xu, Haiying Shen, Chong Ma, Min Shen, Le Chen, Chengzhong Xu, Lin Qu, and Kejiang Ye
    In Proceedings of the 2025 ACM Symposium on Cloud Computing, 2025
  5. ICDCS24
    Quart: Latency-Aware FaaS System for Pipelining Large Model Inference
    Yanying Lin, Yanbo Li, Shijie Peng, Yingfei Tang, Shutian Luo, Haiying Shen, Chengzhong Xu, and Kejiang Ye
    In Proceedings of the 44th IEEE International Conference on Distributed Computing Systems, 2024