Skip to content
View AMark-CS's full-sized avatar

Highlights

  • Pro

Block or report AMark-CS

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AMark-CS/README.md

Hi, I'm Mark πŸ‘‹

MSc in Information Technology @ HKUST Β· Agent Inference Optimization Intern @ Kuaishou

πŸ”¬ Research Focus: Agentic Reasoning Β |Β  AI Agents Β |Β  LLM Post-Training


🧭 About Me

  • πŸŽ“ M.Sc. in Computer Science, The Hong Kong University of Science and Technology (HKUST)
  • 🏒 Algorithm / Engineering Intern @ Kuaishou (快手) β€” working on Agent inference optimization
  • 🧠 Research interests at the intersection of RLHF / GRPO post-training, LLM agents, and efficient inference
  • 🌱 Currently reading: DeepSeek-V4-Pro, Qwen3, GRPO series papers; reproducing SFT β†’ DPO β†’ GRPO pipelines on Qwen
  • πŸ’¬ Reach me on Gmail: hhhzhchhh939@gmail.com

πŸ”¬ Research & Engineering Focus

Area Topics
Post-Training SFT Β· DPO Β· GRPO Β· PPO Β· Online RL Β· Reward Modeling Β· KL-regularized Optimization
AI Agents ReAct Β· Plan-and-Execute Β· Tool Use Β· Function Calling Β· MCP Β· Memory Systems Β· Multi-Agent
Agent Inference KV Cache Compression Β· PagedAttention Β· Speculative Decoding Β· Continuous Batching Β· Quantization (INT8/INT4/AWQ/GPTQ)
Reasoning Models Verifiable Rewards Β· Test-time Scaling Β· Chain-of-Thought Β· Process Reward Models
Distributed Training ZeRO-1/2/3 Β· Tensor Parallel Β· Pipeline Parallel Β· DeepSpeed Β· FSDP

πŸ’» Tech Stack

Languages Python C++ CUDA SQL Bash

ML / Training Frameworks PyTorch Transformers DeepSpeed JAX

Agent / RL / Inference vLLM LangChain LangGraph TRL VERL

Tools & Infra Docker Kubernetes Linux Git LaTeX


πŸ› οΈ What I'm Working On

  • ⚑ Kuaishou Β· Agent Optimization β€” Focusing on ReAct and P&E to improve the inference efficiency of agent.
  • πŸ§ͺ Reproducing Qwen2/3 pipelines β€” SFT + DPO + GRPO end-to-end on single-GPU setups
  • πŸ“– Paper reading group β€” DeepSeek-R1, Qwen3, GRPO/DAPO/GSPO variants, MTP, MLA
  • πŸ“ Drafting a personal notes repo on modern post-training (PPO β†’ DPO β†’ GRPO β†’ Online RLVR)

πŸ“š Open-Source Projects I'm Studying

πŸ€– Agent Frameworks

Repo What I'm learning
HKUDS/OpenHarness Production agent harness design β€” tool use, skills, memory, multi-agent coordination
lsdefine/GenericAgent Self-evolving agent with skill-tree growth from a 3K-line seed β€” a great study of minimalism
langchain-ai/langgraph Graph-based agent orchestration; stateful, long-running agents
All-Hands-AI/OpenHands Code-writing agent architecture; tool calling & sandboxed execution

🎯 Post-Training (RLHF / GRPO)

Repo What I'm learning
verl-project/verl ByteDance's production-grade RL training framework for LLMs (PPO, GRPO, DAPO)
huggingface/trl Reference impl of SFT, DPO, PPO, GRPO on the πŸ€— ecosystem
hiyouga/LLaMA-Factory Unified training for 100+ LLMs β€” SFT, DPO, PPO, ORPO all in one stack
mbzuai-oryx/Awesome-LLM-Post-training Curated survey of post-training papers, code, and benchmarks

⚑ Inference & Systems

Repo What I'm learning
vllm-project/vllm PagedAttention, continuous batching, speculative decoding β€” core of my intern work
sgl-project/sglang RadixAttention & structured generation for LLM agents
NVIDIA/TensorRT-LLM Production-grade inference engine, kernel fusion, in-flight batching

πŸ“– Notes & Interview Prep

Repo What I'm learning
Junvate/LLM-Algorithm-Intern-Guide Hand-derived PPO/RoPE/Transformer notes; DeepSeek & Qwen tech reports
wdndev/llm_interview_note Classic Chinese LLM interview notes β€” distributed, RLHF, alignment
Lau-Jonathan/LLM-Agent-Interview-Guide ByteDance Top-20 high-frequency questions across 10 modules
windsunboyu/post-training-of-llms Chinese translation of DeepLearning.AI's Post-Training course

🐍 GitHub Activity

GitHub Contribution Snake

🀝 Open to

  • πŸ§ͺ Research collaboration on post-training, RLHF, reasoning models
  • πŸ’Ό Full-time opportunities in LLM agent / post-training / inference
  • 🌏 Connecting with researchers & engineers in the HK / SZ / Beijing / Shanghai / Hangzhou AI
  • πŸ“¬ Open-source contributions β€” happy to send PRs to vLLM / VERL / TRL

πŸ“« Contact

Channel Handle
βœ‰οΈ Email hhhzhchhh939l@gmail.com
πŸ™ GitHub @AMark-CS

πŸ€– Building agents that reason, optimizing inference that scales.

Pinned Loading

  1. opencode-buddy opencode-buddy Public

    buddy for opencode

    JavaScript 4

  2. mini-posttrain mini-posttrain Public

    Mini Post-training Repo

    Python 1

  3. hkust-edx-hub hkust-edx-hub Public

    code for hkust edx msit projects

    Python 2

  4. toolglass toolglass Public

    Inspect the mcp tool use by agent

    Python 1

  5. svg-animation-pipeline svg-animation-pipeline Public

    svg-animation-pipline

    TypeScript 1

  6. scientific-agent-skills scientific-agent-skills Public

    Forked from K-Dense-AI/scientific-agent-skills

    Turn any AI agent into an AI Scientist. The #1 Agent Skills library for science, used by 190,000+ scientists worldwide. 165 ready-to-use validated skills plus 100+ scientific databases covering bio…

    Python 1