Yuangang Li

PhD Student at UCI

Yuangang Li
ML Research Intern @ Kilby Labs, Texas Instruments On the Generative AI team, working on coding agents and agentic AI systems for complex task automation.

Hi 👋! I am Yuangang Li, a PhD student at the Donald Bren School of Information and Computer Sciences, University of California, Irvine, where I am fortunate to be advised by Prof. Cristina (Crista) Lopes.

Before my PhD study at UCI, I spent wonderful years at the University of Southern California, completing my Master’s degree and working with Prof. Yue Zhao. During that time, I also had the pleasure of collaborating with Prof. Xiyang Hu and Prof. Jiechao Gao.

My research interests are centered around trustworthy AI, AI agents evaluation and optimization, and AI auditing (programming-language verification for AI).

Research Interests

Trustworthy AI & LLM Safety

Hallucination Mitigation via Causal Reasoning
AAAI 2026
LLM-Based Anomaly Detection & Benchmarking
EMNLP 2025 ACL 2025
Interpretability, Transparency & Social Behavior
ACL 2026 arXiv 2026 arXiv 2025
Reliability in High-Stakes Domains
AAAI 2026 BIBM 2024

Agent Evaluation & Optimization

Reasoning Quality of Coding Agents
arXiv 2026
Benchmarks & Infrastructure for Frontier Agents
arXiv 2026 Harbor-Index Harbor Terminal-Bench 3
Agentic Systems for Real-World Automation
arXiv 2026

Privacy-Preserving & Decentralized Learning

Communication-Efficient Federated Learning
ACM MM 2024 arXiv 2025
Federated Learning for Healthcare & Urban Systems
BIBM 2024 ICDM 2024

News

  • [Jul 2026] Harbor-Index 1.0 is out, a compact, high-signal benchmark for frontier agents: 82 tasks distilled from 6,627 candidates across 54 benchmarks integrated as Harbor adapters, spanning software engineering, research, tool use, mathematics, data analytics, and security. No agent-model pair tested clears 30%. I served as a Core Contributor on the index. [Site] [Contributors]
  • [Jun 2026] New community benchmark: Agents' Last Exam, 1,000+ economically valuable, long-horizon agent tasks contributed by 300+ practitioners and led by UC Berkeley RDI. Frontier agents average 2.6% full pass on the hardest tier. I contributed tasks as one of the authors. [arXiv] [Project]
  • [Jun 2026] Two new preprints with our Stanford collaborators. Clusters are All You Need pre-trains the Tsetlin Machine with semantic clusters drawn from language models, keeping clause-level interpretability while lifting accuracy; The Orchestration Gap examines why process automation stalls in operationally complex industries. [arXiv] [arXiv]
  • [May 2026] I started as a Machine Learning Research Intern at Kilby Labs, Texas Instruments, on the Generative AI team, working on coding agents and agentic AI systems for complex task automation.
  • [Apr 2026] Our paper LLM-Guided Semantic Bootstrapping for Interpretable Text Classification with Tsetlin Machines has been accepted to Findings of ACL 2026! It distills LLM knowledge into interpretable Tsetlin Machines, reaching BERT-level accuracy without embeddings or inference while preserving clause-level transparency. Congratulations to lead author Jiechao Gao and the team. [arXiv] [ACL Anthology]
  • [Apr 2026] New first-author preprint: Beyond Output Correctness. We introduce CodeRQ-Bench, the first benchmark for reasoning quality across code generation, summarization, and classification, analyze 1,069 mismatch cases from existing evaluators, and propose VERA, a two-stage evaluator combining evidence-grounded verification with ambiguity-aware score correction, improving AUCROC by up to 0.26 and AUPRC by up to 0.21. [arXiv] [GitHub]
  • [Jan 2026] I headed to AAAI 2026 in Singapore to present Mitigating Hallucinations in Large Language Models via Causal Reasoning, with S2D-Align also appearing at the conference. Both are now published in the Proceedings of AAAI, Vol. 40. [AAAI] [AAAI]
  • [Dec 2025] I joined the Stanford team behind Terminal-Bench and Harbor as a Research Collaborator, contributing tasks and datasets to Terminal-Bench 3 and adapters to Harbor, the framework for running agent evaluations and building RL environments. [Terminal-Bench] [Harbor]
  • [Nov 2025] Two papers accepted to AAAI 2026! CDCR-SFT internalizes causal DAG-based reasoning during fine-tuning, lifting causal reasoning on CLADDER from 72.9% to 95.3% and cutting hallucinations by 11% on HaluEval; S2D-Align grounds radiology report generation through shallow-to-deep auxiliary alignment. [arXiv] [GitHub] [arXiv]
  • [Sep 2025] I started my PhD at the UC Irvine Donald Bren School of ICS, advised by Prof. Crista Lopes, working on trustworthy LLMs, coding agents, and LLM transparency for code generation.
  • [Aug 2025] Our paper NLP-ADBench has been accepted to Findings of EMNLP 2025! It is the first comprehensive NLP anomaly-detection benchmark, covering 8 datasets and 19 methods, and shows that LLM-based features are consistently strong across the board. [arXiv] [ACL Anthology] [GitHub]
  • [Aug 2025] I wrapped up my Research Assistant positions at Stanford University and USC. Many thanks to all my advisors for their mentorship and support.
  • [May 2025] Our paper AD-LLM: Benchmarking Large Language Models for Anomaly Detection has been accepted to Findings of ACL 2025! It evaluates LLMs for zero-shot detection, data augmentation, and model selection, positioning them as a robust backbone for NLP anomaly detection. Congratulations to lead author Tiankai Yang and the team. [arXiv] [ACL Anthology]
  • [Apr 2025] I started as a Research Assistant at Stanford University, fine-tuning open-source LLMs on the first self-constructed dataset with explicit causal structures and developing LLM-guided semantic bootstrapping for interpretable models.
  • [Dec 2024] New preprint on LLMs for political decision-making: a multi-step reasoning framework that simulates voter decisions, cutting election-prediction error by 77% while exposing bias and overfitting in models like GPT-4o and LLaMA 3.1. [arXiv]
  • [Nov 2024] Two federated-learning papers accepted: FedMetaMed at IEEE BIBM 2024, combining federated and meta-learning for personalized medication across distributed healthcare systems, and Fed-LDR at the IEEE ICDM SSTDM Workshop, using GCNs for spatio-temporal analysis with node-centric refinement. I gave talks on both. [arXiv] [arXiv]
  • [Sep 2024] I began a research collaboration with the University of Virginia on AutoML and LLM inference, working with Dr. Zhaoyuan Su and Prof. Yue Cheng on automated compression strategies for user-specific tasks.
  • [Aug 2024] Co-first-authored review out on Preprints.org: Artificial Intelligence-Aided Digital Twin Design, a systematic survey of how machine learning enhances digital twins across domains. [Preprints]
  • [Jul 2024] Our paper FedBCGD has been accepted to ACM Multimedia 2024! It is the first parameter-block communication method for federated training of large models, letting each client upload a single parameter block per round. [ACM DL] [arXiv]
  • [Jun 2024] New preprint: H-FedSN, a personalized sparse-network method for hierarchical federated learning in IoT that reduces communication cost by up to 238x through structured masking and Bayesian aggregation while preserving accuracy. [arXiv]
  • [Dec 2023] I joined the University of Virginia as a Research Assistant with Dr. Jiechao Gao and Prof. Brad Campbell, working on hierarchical federated learning for IoT and healthcare.
  • [Jul 2023] I joined the USC FORTIS Lab as a Research Assistant, advised by Prof. Yue Zhao, working on trustworthy LLMs and NLP anomaly detection.
  • [Mar 2023] We released Ablator, an open-source deep-learning framework for horizontally scaling ablation experiments with Ray and Optuna, used by 30+ researchers at USC. I also published python-rclone on PyPI, which ships the RClone binary so no pre-installation is needed. [GitHub] [PyPI]
  • [Feb 2023] I joined USC as a Research Engineer with Dr. Iordanis Fostiropoulos, working on distributed AutoML execution and CI infrastructure.
  • [Jan 2023] I joined USC Games as a Game Analyst and Developer on "Hexagon Adventure," building Unity/Firebase telemetry and analyzing player behavior, which drove a 40% increase in engagement.
  • [Dec 2021] I joined SenseTime as an Infrastructure Engineer, building "RocketMQ as a Service" on Kubernetes with an Operator and CRDs that made service creation 50% faster.
  • [Jan 2022] I started my M.S. in Computer Science at the University of Southern California.
  • [Apr 2021] I joined the Institute of Software, Chinese Academy of Sciences as a Research Engineer with Prof. Guoquan Wu, building a record-and-playback web test automation platform that improved end-to-end testing efficiency by 300%.
  • [Jan 2021] I joined NiuTrans as a Back End Developer, building the PDF/XML/EML parsing and translation modules of an AI document translation system that reached 30,000 MAUs.
  • [May 2018] I led development of a recruitment data-mining system at Beijing City University, scraping 10M+ records and building automated dashboards. The project earned National Level Innovative Excellence Project recognition and was adopted by the school.