Spark / Fusion
Performance-oriented Spark platform work, including Fusion-related acceleration and serverless lakehouse analytics.
Alibaba Cloud EMR
Performance-oriented Spark platform work, including Fusion-related acceleration and serverless lakehouse analytics.
OLAP and lakehouse analytics work around EMR StarRocks and Stella Kernel performance leadership.
Vector database work for AI-era retrieval, multimodal data, and vector lake scenarios.
AI-assisted EMR operations and product experience for modern cloud data platforms.
Product and engineering work for cloud data platform development, operations, and user workflows.
Flink, Hive, Spark, StarRocks, Trino, and related open-source engines in Alibaba Cloud EMR.
Hive Metastore, Ranger, and the management layer needed by enterprise data platforms.
Cache, lake format, lake file system, and data lake storage engines for lakehouse and AI-era workloads.
My EMR work has focused on evolving the platform from open-source component management toward cloud-native and serverless data infrastructure, including EMR 2.0, EMR Serverless Spark, lakehouse analytics, and AI-assisted data platform capabilities.
More recently, this direction has expanded into OpenLake, lake-stream integration, multimodal data storage and access, and the shift from human-centric data access patterns to AI-agent-driven data generation and retrieval.
Earlier systems work in HBase, search data infrastructure, realtime compute storage, and Flink state backends shaped the current product direction: production-first, open-source-compatible, and measured under real traffic.