- Contribute across the post-training loop for general-purpose agents: data construction and analysis, evaluation, SFT/RL teacher models, data and model consolidation, and model delivery.
- Work on production-facing Deep Research, adaptive search, and skill-calling behavior, including asynchronous agent workflows.
- Parallel academic explorations led to A²Search and AWA-RL.
Curriculum vitae · July 2026
Fengji Zhang
PhD candidate · Post-training for general-purpose LLM agents
Research profile
PhD candidate in Computer Science at City University of Hong Kong, expecting to graduate in Fall 2027. I work on post-training for general-purpose LLM agents across data, evaluation, reinforcement learning, model behavior, and delivery.
At Alibaba Qwen, I contribute to post-training for Deep Research, adaptive search, and skill use. Earlier work spans reliable search agents, multimodal evaluation, repository context, and execution-guided code generation.
Research experience
- Worked on Yi-VL post-training and evaluation for instruction following and visual understanding.
- Developed and evaluated Yi-Coder for code generation, completion, and coding-agent settings; studied test-time scaling with long reasoning and execution feedback.
- Developed HumanEval-V during this internship as a separate academic project.
- Proposed CodeT, using generated tests and execution agreement to improve inference-time selection for code generation.
- Developed RepoCoder, an iterative retrieval-generation approach for repository-level code completion and cross-file context.
- Developed AIOps and internal SaaS tooling for online-game operations.
Selected papers
To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning
Fengji Zhang, Tianyu Fan, Yuxiang Zheng, Xinyao Niu, Chengen Huang, Jacky Keung, Bei Chen.
AWA-RL learns a controllable answer–abstain policy from query-specific prior capability and on-policy observations.
A²Search: Ambiguity-Aware Question Answering with Reinforcement Learning
Fengji Zhang, Xinyao Niu, Chengyang Ying, et al.
Annotation-free ambiguity resolution with trajectory sampling, evidence verification, and an AnsF1 reward for multiple valid answers.
HumanEval-V: Systematic Evaluation of Visual Reasoning in Large Multimodal Models for Code Generation
Fengji Zhang, Linquan Wu, Huiyu Bai, et al.
A human-annotated benchmark for testing visual understanding through executable coding tasks with indispensable diagram context.
RepoCoder: Repository-level Code Completion through Iterative Retrieval and Generation
Fengji Zhang, Bei Chen, Yue Zhang, et al.
Iterative retrieval-generation for cross-file context; more than 10% improvement over in-file completion baselines across RepoBench settings.
CodeT: Code Generation with Generated Tests
Bei Chen*, Fengji Zhang*, Anh Nguyen*, et al.
Generated tests and dual execution agreement for solution selection; 65.8% pass@1 on HumanEval.
Additional work spans code intelligence, multimodal understanding, deep research, software engineering, and responsible code generation. View the complete publication record on Google Scholar ↗
Education
City University of Hong Kong
PhD candidate in Computer Science · 2023–2027 expected
Wuhan University
MS in Computer Science · 2020–2023
BS in Computer Science · 2016–2020
Service & recognition
Reviewer
ICSE ’24, ICLR ’25/’26, ICML ’26, NeurIPS ’26, CVPR ’25, ICCV ’25, ACL ’25/’26, EMNLP ’26
Award
CityUHK Outstanding Academic Performance