MSc in Information Technology @ HKUST Β· Agent Inference Optimization Intern @ Kuaishou
π¬ Research Focus: Agentic Reasoning Β |Β AI Agents Β |Β LLM Post-Training
- π M.Sc. in Computer Science, The Hong Kong University of Science and Technology (HKUST)
- π’ Algorithm / Engineering Intern @ Kuaishou (εΏ«ζ) β working on Agent inference optimization
- π§ Research interests at the intersection of RLHF / GRPO post-training, LLM agents, and efficient inference
- π± Currently reading: DeepSeek-V4-Pro, Qwen3, GRPO series papers; reproducing SFT β DPO β GRPO pipelines on Qwen
- π¬ Reach me on Gmail: hhhzhchhh939@gmail.com
| Area | Topics |
|---|---|
| Post-Training | SFT Β· DPO Β· GRPO Β· PPO Β· Online RL Β· Reward Modeling Β· KL-regularized Optimization |
| AI Agents | ReAct Β· Plan-and-Execute Β· Tool Use Β· Function Calling Β· MCP Β· Memory Systems Β· Multi-Agent |
| Agent Inference | KV Cache Compression Β· PagedAttention Β· Speculative Decoding Β· Continuous Batching Β· Quantization (INT8/INT4/AWQ/GPTQ) |
| Reasoning Models | Verifiable Rewards Β· Test-time Scaling Β· Chain-of-Thought Β· Process Reward Models |
| Distributed Training | ZeRO-1/2/3 Β· Tensor Parallel Β· Pipeline Parallel Β· DeepSpeed Β· FSDP |
- β‘ Kuaishou Β· Agent Optimization β Focusing on ReAct and P&E to improve the inference efficiency of agent.
- π§ͺ Reproducing Qwen2/3 pipelines β SFT + DPO + GRPO end-to-end on single-GPU setups
- π Paper reading group β DeepSeek-R1, Qwen3, GRPO/DAPO/GSPO variants, MTP, MLA
- π Drafting a personal notes repo on modern post-training (PPO β DPO β GRPO β Online RLVR)
| Repo | What I'm learning |
|---|---|
| HKUDS/OpenHarness | Production agent harness design β tool use, skills, memory, multi-agent coordination |
| lsdefine/GenericAgent | Self-evolving agent with skill-tree growth from a 3K-line seed β a great study of minimalism |
| langchain-ai/langgraph | Graph-based agent orchestration; stateful, long-running agents |
| All-Hands-AI/OpenHands | Code-writing agent architecture; tool calling & sandboxed execution |
| Repo | What I'm learning |
|---|---|
| verl-project/verl | ByteDance's production-grade RL training framework for LLMs (PPO, GRPO, DAPO) |
| huggingface/trl | Reference impl of SFT, DPO, PPO, GRPO on the π€ ecosystem |
| hiyouga/LLaMA-Factory | Unified training for 100+ LLMs β SFT, DPO, PPO, ORPO all in one stack |
| mbzuai-oryx/Awesome-LLM-Post-training | Curated survey of post-training papers, code, and benchmarks |
| Repo | What I'm learning |
|---|---|
| vllm-project/vllm | PagedAttention, continuous batching, speculative decoding β core of my intern work |
| sgl-project/sglang | RadixAttention & structured generation for LLM agents |
| NVIDIA/TensorRT-LLM | Production-grade inference engine, kernel fusion, in-flight batching |
| Repo | What I'm learning |
|---|---|
| Junvate/LLM-Algorithm-Intern-Guide | Hand-derived PPO/RoPE/Transformer notes; DeepSeek & Qwen tech reports |
| wdndev/llm_interview_note | Classic Chinese LLM interview notes β distributed, RLHF, alignment |
| Lau-Jonathan/LLM-Agent-Interview-Guide | ByteDance Top-20 high-frequency questions across 10 modules |
| windsunboyu/post-training-of-llms | Chinese translation of DeepLearning.AI's Post-Training course |
- π§ͺ Research collaboration on post-training, RLHF, reasoning models
- πΌ Full-time opportunities in LLM agent / post-training / inference
- π Connecting with researchers & engineers in the HK / SZ / Beijing / Shanghai / Hangzhou AI
- π¬ Open-source contributions β happy to send PRs to vLLM / VERL / TRL
| Channel | Handle |
|---|---|
| βοΈ Email | hhhzhchhh939l@gmail.com |
| π GitHub | @AMark-CS |
π€ Building agents that reason, optimizing inference that scales.


