Agentic AILarge Language ModelsReinforcement LearningRecursive Self-Improvement
I make AI agents work in the real world:
ClawBench for live web-agent evaluation
(adopted by ByteDance Seed, EMNLP 2026 Findings),
RewardHarness for self-evolving agentic post-training (COLM 2026), and
VidGround for grounded video understanding.
I'm open to collaborations and research chats. Reach out at
reacher [at] cs.ubc.ca ( WeChat works too).
Actively looking for research internships. Let's grab a coffee and have a chat!
Research Highlights
EMNLP 2026 Findings
ClawBench: Can AI Agents Complete Everyday Online Tasks?
ClawBench evaluates AI agents on 153 real-world tasks across 144 live platforms, from booking appointments to filing job applications. It runs on production websites and intercepts only the final submission, keeping evaluation safe without losing real-world complexity.
Language models can answer video questions from text priors alone, without watching the video. We filter out such spurious training samples, producing a cleaner dataset that teaches models to ground answers in what they see.
RewardHarness beats GPT-5 by 5.3 points on image-editing evaluation while using 0.05% of the EditReward data. It reframes reward modeling as context evolution: from 100 preference demonstrations it evolves a library of tools and skills.
We are organizing ARRSI 2026, the first workshop on Autonomous Research and Recursive Self-Improvement, co-located with AACL-IJCNLP 2026. Call for papers is open!
ClawBench has been accepted at the COLM 2026 Workshop on Agent Behavior. See the paper.
ClawBench has been adopted by ByteDance Seed (Doubao).
@article{zhang2026rewardharness,
title={RewardHarness: Self-Evolving Agentic Post-Training},
author={Zhang, Yuxuan and Du, Penghui and Li, Bo and Wei, Cong and Miao, Junwen and Zhang, Huaisong and Cai, Songcheng and Wang, Yubo and Jiang, Dongfu and Zhang, Yuyu and others},
journal={arXiv preprint arXiv:2605.08703},
year={2026}
}
★ClawBench: Can AI Agents Complete Everyday Online Tasks?
Yuxuan Zhang, Yubo Wang,
Yipeng Zhu,
Penghui Du,
Junwen Miao,
Xuan Lu,
Zhuofeng Li,
Xingwei Qu,
Zhengkang Guo,
Yuanzhe Shen,
Dingjie Song,
Han Zhou,
Tuney Zheng,
Xian Wu,
Hao Yu,
Songcheng Cai,
Yi Lu,
Yunzhuo Hao,
Minyi Lei,
Liang Chen,
Kai Zou,
Huifeng Yin,
Wendong Xu, Dongfu Jiang,
Ping Nie,
Jiaheng Liu, Wenhu Chen, Kelsey R. Allen
@article{zhang2026clawbench,
title={ClawBench: Can AI Agents Complete Everyday Online Tasks?},
author={Zhang, Yuxuan and Wang, Yubo and Zhu, Yipeng and Du, Penghui and Miao, Junwen and Lu, Xuan and Li, Zhuofeng and Qu, Xingwei and Guo, Zhengkang and Shen, Yuanzhe and others},
journal={arXiv preprint arXiv:2604.08523},
year={2026}
}
★Watch Before You Answer: Learning from Visually Grounded Post-Training
@article{zhang2026watch,
title={Watch before you answer: Learning from visually grounded post-training},
author={Zhang, Yuxuan and Hwang, EunJeong and Zhang, Huaisong and Du, Penghui and Jia, Yiming and Jiang, Dongfu and He, Xuan and Zhang, Shenhui and Nie, Ping and West, Peter and others},
journal={arXiv preprint arXiv:2604.05117},
year={2026}
}
WebWorld: The Browser as a World Model for Self-Improving Web Code
Jiajun Wu,
Jian Yang,
Yaxin Du,
Wei Zhang,
Haowen Wang,
Junhang Cheng, Yuxuan Zhang,
Tuney Zheng,
Xianglong Liu,
Ming Zhou
EMNLP 2026
@inproceedings{wu2026webworld,
title={WebWorld: The Browser as a World Model for Self-Improving Web Code},
author={Wu, Jiajun and Yang, Jian and Du, Yaxin and Zhang, Wei and Wang, Haowen and Cheng, Junhang and Zhang, Yuxuan and Zheng, Tuney and Liu, Xianglong and Zhou, Ming},
booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year={2026}
}
★VGI-BENCH: Probing Visual Intelligence in Video Generation Models
Xuan He†§,
Cong Wei†,
Yuhao Cheng†,
Linrui Ma†, Yuxuan Zhang†,
Zuojun Li,
Yuhao Wen,
Zeyi Liu,
Yuren Hao,
Songcheng Cai,
Keming Wu,
Penghui Du,
Kai Zou,
Rui Yang,
Chenkai Sun,
Ke Yang,
Ping Nie,
Kelsey R Allen,
Chenglong Wang,
Michel Galley,
Jianfeng Gao,
ChengXiang Zhai(†Main Contributor. §Project Lead.)
@inproceedings{he2026vgibench,
title={VGI-BENCH: Probing Visual Intelligence in Video Generation Models},
author={He, Xuan and Wei, Cong and Cheng, Yuhao and Ma, Linrui and Zhang, Yuxuan and Li, Zuojun and Wen, Yuhao and Liu, Zeyi and Hao, Yuren and Cai, Songcheng and Wu, Keming and Du, Penghui and Zou, Kai and Yang, Rui and Sun, Chenkai and Yang, Ke and Nie, Ping and Allen, Kelsey R. and Wang, Chenglong and Galley, Michel and Gao, Jianfeng and Zhai, ChengXiang},
booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year={2026}
}
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
@article{wang2026function,
title={Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models},
author={Wang, Yubo and Liang, Jiarong and Zhang, Yuxuan and Liu, Xuye and Wei, Cong and Zhang, Yuyu and Nie, Ping and Chen, Wenhu},
journal={arXiv preprint arXiv:2607.12463},
year={2026}
}
OpenSkill: Open-World Self-Evolution for LLM Agents
Zhiling Yan,
Dingjie Song,
Hanrong Zhang,
Wei Liang, Yuxuan Zhang,
Yutong Dai,
Lifang He,
Philip S. Yu,
Ran Xu,
Xiang Li,
Lichao Sun
@article{yan2026openskill,
title={OpenSkill: Open-World Self-Evolution for LLM Agents},
author={Yan, Zhiling and Song, Dingjie and Zhang, Hanrong and Liang, Wei and Zhang, Yuxuan and Dai, Yutong and He, Lifang and Yu, Philip S and Xu, Ran and Li, Xiang and others},
journal={arXiv preprint arXiv:2606.06741},
year={2026}
}
MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection
Haowen Wang,
Yaxin Du,
Jian Yang,
Jiajun Wu,
Shukai Liu, Yuxuan Zhang,
Pingjie Wang,
Siheng Chen,
Tuney Zheng,
Ming Zhou,
Xianglong Liu,
Bryan Dai
@article{wang2026mira,
title={MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection},
author={Wang, Haowen and Du, Yaxin and Yang, Jian and Wu, Jiajun and Liu, Shukai and Zhang, Yuxuan and Wang, Pingjie and Chen, Siheng and Zheng, Tuney and Zhou, Ming and others},
journal={arXiv preprint arXiv:2605.30288},
year={2026}
}
Dr. Claw: A Unified System for the Vibe Research Paradigm
Dingjie Song,
Hanrong Zhang,
Dawei Liu,
Yixin Liu,
Zongxia Li,
Zhengqing Yuan,
Siqi Zhang,
Henry Peng Zou,
Zhiling Yan, Yuxuan Zhang,
Yanfang Ye,
Philip S. Yu,
Lichao Sun
@inproceedings{song2026drclaw,
title={Dr. Claw: A Unified System for the Vibe Research Paradigm},
author={Song, Dingjie and Zhang, Hanrong and Liu, Dawei and Liu, Yixin and Li, Zongxia and Yuan, Zhengqing and Zhang, Siqi and Zou, Henry Peng and Yan, Zhiling and Zhang, Yuxuan and Ye, Yanfang and Yu, Philip S. and Sun, Lichao},
booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing: System Demonstrations},
year={2026},
url={https://github.com/OpenLAIR/dr-claw}
}
20255 papers
Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
Yang Yao,
Yixu Wang, Yuxuan Zhang,
Yi Lu,
Tianle Gu,
Lingyu Li,
Dingyi Zhao,
Keming Wu, Haozhe Wang,
Ping Nie,
Yan Teng,
Yingchun Wang
@article{yao2026dr,
title={Dr. bench: A multidimensional evaluation for deep research agents, from answers to reports},
author={Yao, Yang and Wang, Yixu and Zhang, Yuxuan and Lu, Yi and Gu, Tianle and Li, Lingyu and Zhao, Dingyi and Wu, Keming and Wang, Haozhe and Nie, Ping and others},
year={2026}
}
ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations
@article{wang2025scholarcopilot,
title={Scholarcopilot: Training large language models for academic writing with accurate citations},
author={Wang, Yubo and Ma, Xueguang and Nie, Ping and Zeng, Huaye and Lyu, Zhiheng and Zhang, Yuxuan and Schneider, Benjamin and Lu, Yi and Yue, Xiang and Chen, Wenhu},
journal={arXiv preprint arXiv:2504.00824},
year={2025}
}
Retri3D: 3D Neural Graphics Representation Retrieval
@inproceedings{guan2025retri3d,
title={Retri3D: 3D Neural Graphics Representation Retrieval},
author={Guan, Yushi and Kwan, Daniel and Dandurand, Jean and Yan, Xi and Liang, Ruofan and Zhang, Yuxuan and Jain, Nilesh and Ahuja, Nilesh and Panneer, Selvakumar and Vijaykumar, Nandita},
booktitle={International Conference on Learning Representations},
volume={2025},
pages={31771--31813},
year={2025}
}
VideoScore2: Think before You Score in Generative Video Evaluation
@article{he2025videoscore2,
title={Videoscore2: Think before you score in generative video evaluation},
author={He, Xuan and Jiang, Dongfu and Nie, Ping and Liu, Minghao and Jiang, Zhengxuan and Su, Mingyi and Ma, Wentao and Lin, Junru and Ye, Chun and Lu, Yi and others},
journal={arXiv preprint arXiv:2509.22799},
year={2025}
}
StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs
@article{yang2025structeval,
title={StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs},
author={Yang, Jialin and Jiang, Dongfu and He, Lipeng and Siu, Sherman and Zhang, Yuxuan and Liao, Disen and Li, Zhuofeng and Zeng, Huaye and Jia, Yiming and Wang, Haozhe and others},
journal={arXiv preprint arXiv:2505.20139},
year={2025}
}
AdoptedClawBench is used as a headline agent benchmark in frontier model technical reports. Li Auto's Mach-Mind-4-Flash reports its ClawBench score in the abstract and benchmarks eight model families on it, including Kimi K2.5, Qwen3.5, GLM-4.7 and Nemotron-3. 2026
CitedClawBench has been cited by agent-benchmark work from Shanghai AI Laboratory (WildClawBench, π-Bench), Carnegie Mellon (MyPCBench), HKU MMLab and Meituan (UniClawBench), and Peking University (Harness-Bench). 2026
IntegratedVideoScore2 is used as a reward model by Google's VQQA video-evaluation framework, which runs it as a Best-of-N selector across its main results tables. 2026
CitedScholarCopilot is surveyed in "Transforming Science with Large Language Models," a widely cited review of AI-assisted scientific discovery. 2025 - 2026