Senior Software Development Engineer · Guangzhou, China
Ekko: low-latency model update for multi-terabyte DLRMs (published in part as OSDI '22)2.4 s model-update latency and 10,000× model-size scaling for multi-terabyte recommendation models; serves 1 B+ users daily in WeChat.
- Problem. Scaling DLRMs improved offline accuracy but degraded online engagement; root cause: stale models from increased model-update latency.
- Key idea. Co-designed deployment mechanisms with model-aware policies (compressed update dissemination, accuracy-aware scheduling, SLO-aware placement, safe rollback).
- Technical contributions. WAN bandwidth −92 %, machine cost −49 %, 2.4 s model-update latency; 10,000× model-size scaling (GB → tens of TB).
- Outcomes. Core techniques published as OSDI '22 (co-first author). Deployed in WeChat recommendation stacks, serves 1 B+ users daily. Official WeChat blog reports +40 % DAU and +87 % total VV over six months after full adoption (alongside product iteration and operations).
Data and feature platform: safe, scalable pipelinesWebAssembly-based runtime with in-process isolation; data movement reduced up to 1,200× on representative workloads.
- Problem. Modern feature pipelines are long and increasingly multimodal; cross-process operator composition creates high overhead and expensive data movement.
- Approach. WebAssembly-based runtime for in-process isolation (safety + resource constraints) and locality-aware operator placement near data sources.
- Outcome. Data movement reduced up to 1,200× on representative workloads; widely used within WeChat for data preparation.
