Yenkai Huang

Yenkai Huang — AI Platform PM / Agent Infrastructure PM / Developer Platform PM / Kubernetes Platform PM / GenAI Product PM

I build the infrastructure AI agents actually run on: agent runtimes, sandbox isolation, and the developer platform underneath them.

Work

Eight problems I've owned across agent infrastructure, developer platforms, enterprise AI, analytics, and governance. Open any of them for the full story.

Agent Runtime & Isolation

Running untrusted agents in production

An agent that executes code is an untrusted workload holding your credentials. Most teams solve that by not letting agents do anything interesting. We built real isolation instead.

Runs in production behind a customer-facing AI product.
gVisorCilium eBPFKubernetes
Read →
Developer Platform

Platform operations as agent-executable skills

Platform teams write runbooks. Engineers don't read them — they file a ticket or they guess. The knowledge exists; the interface is wrong.

~10k users can move from stated intent to a governed production action.
MCPDeveloper experienceGovernance
Read →
AI Operations

Rules for the 80%, models for the tail

Most AI-for-incidents demos showcase the novel failure. Production is the opposite: the same signatures recur, and at 3am you want deterministic, not creative.

Triages incidents in seconds, with 20%+ classification accuracy.
Knowledge graphMulti-agentAIOps
Read →
Cost & Capacity

Right-sizing a fleet with a seasonal load curve

Generic cost optimizers assume steady-state traffic. A business with a hard seasonal peak has the opposite requirement.

1,500+ services; 15%+ cost reduction; 1,000+ engineering hours saved.
KarpenterCRD controllersFinOps
Read →
Enterprise GenAI

Shipping generative analytics before the playbook existed

A useful enterprise AI product is as much a legal, compliance, and workflow problem as a model problem. The product had to make natural-language analysis trustworthy enough to use.

Reduced quarterly business-review reporting time by 70%.
Text2SQLPythonEnterprise AI
Read →
RAG Product

Grounding a sales assistant in renewal knowledge

Sales teams needed answers assembled from enterprise knowledge, not another chatbot that improvised. Retrieval and governance were the product foundations.

Delivered a GPT assistant for sales-renewal workflows.
RAGWebexData governance
Read →
AIOps & Observability

Building the operational view from zero

Managing infrastructure as SaaS requires a shared, real-time view of system behavior. Without it, every support case begins with reconstructing reality.

Shipped SaaS manageability to 2,000+ enterprise accounts.
KafkaSnowflakeTableau
Read →
Responsible AI

Scaling content governance with policy and data

At platform scale, moderation cannot be a collection of local judgment calls. Policy, escalation data, and model quality have to operate as one system.

Improved AI classification accuracy by 10%.
AI governancePolicyClassification
Read →