Field Application Engineer @ Taiwan AILabs
I deploy on-premises LLM systems into customer environments and contribute fixes across the GPU-inference stack. My work spans Linux/Kubernetes operations, GPU serving, customer-facing troubleshooting, CUDA kernels, and early Blackwell (SM120/NVFP4) enablement.
Portfolio · PR wall · LinkedIn · CV
- Customer deployment: turn site constraints into deployment plans, acceptance criteria, and production handover for on-premises LLM systems.
- GPU inference: benchmark and debug serving, communication, quantization, and kernel paths across vLLM, TensorRT-LLM, SGLang, Triton, Dynamo, and FlashInfer.
- Upstream correctness: 17 merged/landed changes, with the complete live record on prs.wayne.is-a.dev.
| Project | Evidence |
|---|---|
| trtllm-triton-serving | TensorRT-LLM vs vLLM on H100; 12 controlled studies |
| tensor-core-from-scratch | 10 CUDA matmul kernels from naive to Tensor Cores |
| inference-kernel-cookbook | Flash Attention, KV cache, and paged attention from scratch |
| nccl-collectives-bench | NCCL bandwidth, latency, NVLS, and TP-decode limits |
| nim-agent-blueprint | NVIDIA NIM agentic RAG with evaluation and observability |
| llm-security-lab | Reproducible LLM attacks and defenses |



