About Me

Hi there! I am Yuanqing Wang (王源清), and I also go by Yuan. I am a Mathematics PhD student at The University of Texas at Dallas, advised by Professor Baris Coskunuzer.

My research focuses on post-training for long-horizon LLM agents, particularly agentic reinforcement learning, credit assignment, and self-improving agents.

Before coming to UT Dallas, I received my M.S. in Mathematics from Capital Normal University, advised by Professor Zhenlei Zhang, and my B.S. in Information and Computational Science from Xiamen University.

I am actively seeking internship opportunities and would be happy to connect with researchers and teams working on related problems. Please feel free to reach out!

News

  • 2026.08 Co-first-author paper dwDPO accepted to EMNLP 2026 Main, and Skill-CDPO accepted to EMNLP 2026 Findings.

Publications

dwDPO teaser Image pending

dwDPO: Training Multi-Turn LLM Agents via Divergence-Weighted Direct Preference Optimization

Jinghao Lin*, Yuhang Wu*, Yuanqing Wang*, Yuchen Li, Kangtianxingjian, Baris Coskunuzer, Xiawu Zheng

EMNLP 2026 Main

Uses implicit rewards to pinpoint the critical steps in long-horizon agent trajectories, then weights each step by a principled measure of how much it mattered. The weighting is cheap and drop-in, and gives stable gains over standard DPO across six benchmarks.

Skill-CDPO teaser Image pending

Skill-CDPO: Evolving Agent Tool-Use via Critical Step Preference Optimization

Yuchen Li, Jinghao Lin, Yuanqing Wang

EMNLP 2026 Findings

Finds where an agent's tool use breaks down by comparing expert and local rollouts, then builds preference pairs weighted by step criticality and score gap. An 8B model trained this way matches or beats GPT-5.2 on medical agent benchmarks.

Heat Field Signatures teaser Image pending

Heat Field Signatures: From Point Clouds to Smooth Geometry

Yuanqing Wang, Yapeng Tian, Baris Coskunuzer

Uses differential geometry and training-free heat-field features for point-cloud classification. HFS variants lead seven main benchmarks; HFS-full achieves 59.6% versus 35.7% mean class accuracy for Point Transformer on SCOP, with an 8.5× median end-to-end speedup over DGCNN across four datasets.

Multimodal Learning and Evaluation

First author

In Submission, ACL ARR

Research on multimodal understanding and model evaluation.

Education

  • 2024.08 - Present, The University of Texas at Dallas, Richardson, TX
    PhD Student in Mathematics

  • 2021.09 - 2024.06, Capital Normal University, Beijing, China
    M.S. in Mathematics
    Master’s thesis: Notes on Calabi Conjecture and Kähler–Einstein Metric — English · 中文

  • 2016.09 - 2021.06, Xiamen University, Xiamen, China
    B.S. in Information and Computational Science (Elite Students’ Program)

Experience

  • 2021.06 - 2021.08, Gathssen Investment Co., Ltd., Shenzhen, China
    Quantitative Strategy Developer Intern