Skip to content
View xu-wentao's full-sized avatar
📖
📖
  • Beijing.China

Block or report xu-wentao

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Xu-Wentao/README.md

Hi, I'm Wentao Xu

Typing SVG

Cloud-native and AI infrastructure engineer working across Kubernetes, heterogeneous accelerator scheduling, observability, distributed systems, and LLM infrastructure.

I spend much of my engineering time studying production problems, turning them into reusable designs, and contributing the resulting improvements back to open-source communities.

Open-source focus

  • Kubernetes scheduling and resource management for AI workloads
  • Dynamic Resource Allocation, GPU / NPU devices, and multi-tenant quota control
  • Operators, controllers, observability, and production reliability
  • Cloud-native database automation and disaster recovery
  • Go-based infrastructure components and platform tooling

Community work

  • Volcano — contributing to scheduler capabilities around Kubernetes DRA, accelerator-aware quota management, queue semantics, and control-plane efficiency.
  • Prometheus Operator — working on operator-managed authentication, workload generation, probes, and production-oriented observability configuration.
  • Apache ShardingSphere on Cloud — contributed to cloud database lifecycle automation, PITR and backup workflows, reliability fixes, and operator tooling.
  • NVIDIA GPU Operator — contributed fixes around monitoring integration and GPU observability configuration.

My contributions usually start from operational problems seen in real clusters: resource isolation, backward compatibility, observability, recovery workflows, or making complex platform behavior easier to operate.

What I care about

Infrastructure should remain understandable and dependable, even when the systems running on it become increasingly complex.

I am particularly interested in work where scheduler design, Kubernetes APIs, accelerator topology, observability, and operational experience meet.

Engineering toolkit

Cloud Native & Platform

Cloud Native

Languages & Development

Languages

Delivery & Automation

Tooling

Kubernetes Scheduling · DRA · Volcano · GPU / NPU · LLM Serving · Operators · SRE


中文简介:云原生与 AI 基础设施工程师,持续参与 Volcano、Prometheus Operator、Apache ShardingSphere on Cloud 等开源社区,关注 Kubernetes 调度、DRA、异构算力管理、可观测性以及大模型训练与推理基础设施。

Pinned Loading

  1. apache/shardingsphere-on-cloud apache/shardingsphere-on-cloud Public

    A collection of tools and best practices to take ShardingSphere into the cloud

    Go 88 33

  2. prometheus-operator/prometheus-operator prometheus-operator/prometheus-operator Public

    Prometheus Operator creates/configures/manages Prometheus clusters atop Kubernetes

    Go 10k 3.9k

  3. volcano-sh/volcano volcano-sh/volcano Public

    A Cloud Native Batch System (Project under CNCF)

    Go 5.8k 1.5k