Cloud-native and AI infrastructure engineer working across Kubernetes, heterogeneous accelerator scheduling, observability, distributed systems, and LLM infrastructure.
I spend much of my engineering time studying production problems, turning them into reusable designs, and contributing the resulting improvements back to open-source communities.
- Kubernetes scheduling and resource management for AI workloads
- Dynamic Resource Allocation, GPU / NPU devices, and multi-tenant quota control
- Operators, controllers, observability, and production reliability
- Cloud-native database automation and disaster recovery
- Go-based infrastructure components and platform tooling
- Volcano — contributing to scheduler capabilities around Kubernetes DRA, accelerator-aware quota management, queue semantics, and control-plane efficiency.
- Prometheus Operator — working on operator-managed authentication, workload generation, probes, and production-oriented observability configuration.
- Apache ShardingSphere on Cloud — contributed to cloud database lifecycle automation, PITR and backup workflows, reliability fixes, and operator tooling.
- NVIDIA GPU Operator — contributed fixes around monitoring integration and GPU observability configuration.
My contributions usually start from operational problems seen in real clusters: resource isolation, backward compatibility, observability, recovery workflows, or making complex platform behavior easier to operate.
Infrastructure should remain understandable and dependable, even when the systems running on it become increasingly complex.
I am particularly interested in work where scheduler design, Kubernetes APIs, accelerator topology, observability, and operational experience meet.
Kubernetes Scheduling ·
DRA ·
Volcano ·
GPU / NPU ·
LLM Serving ·
Operators ·
SRE
中文简介:云原生与 AI 基础设施工程师,持续参与 Volcano、Prometheus Operator、Apache ShardingSphere on Cloud 等开源社区,关注 Kubernetes 调度、DRA、异构算力管理、可观测性以及大模型训练与推理基础设施。


