Open, unified data infrastructure for the AI era

Building cloud-scale data infrastructure.

I am Yu Li, EMR Lead at Alibaba Cloud and a long-time Apache contributor. My work focuses on distributed computing, cloud-native big-data platforms, lakehouse architecture, stream processing, and data infrastructure for AI.

About

At Alibaba Cloud, I lead the E-MapReduce (EMR) team. My role combines product direction, engineering execution, and open-source collaboration across cloud data platforms.

I am an ASF Member and serve in the Apache ecosystem across Flink, HBase, Paimon, Celeborn, Gluten, Fluss, Amoro, and GraphAr. My Apache work connects production requirements with community-governed infrastructure across storage, state management, lakehouse systems, and AI-era data platforms.

Earlier at Alibaba, I worked on computing platform and search infrastructure, including HBase, large-scale cluster operations, Flink state storage, and open-source community work in China. Before Alibaba, I worked on IBM InfoSphere BigInsights after completing graduate studies at Beihang University.

Background

Current Role

EMR Lead at Alibaba Cloud, responsible for cloud data platform products and engineering teams. I joined Alibaba in 2013.

Education

Beihang University. M.S. in Computer System Architecture, 2007-2010; B.S. in Computer Science and Technology, 2003-2007.

Expertise

Distributed computing, open-source community development, cloud data platforms, and large-scale systems engineering.

Career Timeline

Since completing graduate studies, I have worked at two companies: IBM from 2010 to 2013, and Alibaba from 2013 to today. The timeline below shows how my responsibilities and scope have grown along that path.

  • Alibaba Cloud: Head of EMR, covering DataOps, open-source computing engines, metadata and security, and lake storage systems.
  • Alibaba Cloud: Led Flink storage and EMR platform teams, including checkpoint, savepoint, state backend, EMR console, monitoring, studio, and resilience work.
  • Alibaba Cloud: Led real-time compute storage work, focusing on Flink state backend performance, stability, and production requirements.
  • Alibaba Search: Technical lead of the HBase team, including high-throughput storage work that reached 100K QPS per node during Double 11 production traffic.
  • Alibaba Search: Core member of the data platform team, working on Hadoop/HBase clusters serving hybrid index-building and real-time query workloads at large scale.
  • Alibaba: Member of the HBase team, developing and maintaining HBase for business requirements and big-data production workloads.
  • IBM China Software Development Laboratory: Staff Software Engineer, working on InfoSphere BigInsights and leading open-source team efforts.

Academic Work

Before my long-running work in Apache and cloud-scale data infrastructure, my academic work focused on heterogeneous cluster scheduling and web information retrieval. Google Scholar and ORCID 0000-0003-4355-2764 list those early publications.

Selected Papers

Recent Highlights

  • Flink Forward Asia talk on building an AI-native multimodal lake with Apache Paimon and Milvus.
  • Yunqi Conference sessions on EMR AI, Fusion/Stella TPC performance leadership, and Alibaba Cloud Milvus.
  • DataFun Data+AI Summit talk on OpenLake, a lakehouse platform solution for the AI era.
  • Alibaba Cloud EMR Serverless Spark commercialization and lakehouse analytics talks highlighted the EMR serverless transition.
  • Shared Alibaba Cloud EMR lakehouse analytics work with the StarRocks community.
  • Continued work on Flink state storage and snapshots while contributing to HBaseCon Asia community activities.