News iconNews

Academic

news icon One first-author paper accepted by NeurIPS 2026: State of Thought.

announcement icon We released SoT (a new endogenous reasoning paradigm) — that's an interesting work, let's share our insight!

news icon One paper accepted by IEEE WCL: Channel Gains to Captions.

announcement icon We released CoCurve ((M)LLM pruning via 2nd-order curvature) — let's talk!

announcement icon We released CoAx (mechanistic interpretability of Transformers) — welcome to reach out and discuss!

news icon Two first-author papers accepted by ICML 2026: SubspacePath Pruner and XDomainBench.

news icon SubspacePath Pruner and Rule-Guided Active Inference accepted by ICML AdaptFM Workshop and ICML FoGen Workshop.

news icon One first-author paper accepted by ICLR 2026: Rule-Guided Active Inference.

Personal

Joining MSRA as a Research Intern.

Joined NTU as a PhD candidate and received the NTU Research Scholarship.

Joined CUHKSZ as a Research Assistant.

☕ Let's coffee chat

I stay critical of my own research and am always looking for sharper, more insightful, and more forward-looking ideas. If that resonates with you, feel free to reach out anytime. I'd love to chat over coffee.

Research map iconResearch Map

I investigate the capability frontier of foundation models under efficient, scalable, and reliable real-world deployment. My research exploits the interpretable structural information within (M)LLMs and agentic systems to advance both intrinsic model efficiency and inference-time reasoning performance, turning a mechanistic understanding of their internal computation into trustworthy deployment at scale.

RQ1Foundation Model Internal Structure

How a frozen model computes inside, from circuits to experts, so later gains rest on mechanism, not guesswork.

State of Thought1stNeurIPS’26↗
SubspacePath Pruner1stICML’26↗
CoAx1stPreprint↗
RLTBSoon
MoE Expert & Sample TheorySoon

RQ3Reasoning Efficiency

Route internal pathways, pick the evidence each step needs, and stop early, cutting tokens and latency at run time.

SubspacePath Pruner1stICML’26↗
CoCurve1stPreprint↗
Ada Merging & Distillation1stSoon↗
SkillFocusPreprint
SpecRoute Agent ServingSoon
State of Thought1stNeurIPS’26↗
Active Inference1stICLR’26↗
LoTR1stSoon↗
Cross-Agent MemorySoon
ParetoEvoSoon

RQ2Model / Agent Efficiency

Pruning, co-pruning, budget-aware merging, and faster serving make the model and agent cheap before deployment.

XDomainBench1stICML’26↗
Channel GainsIEEE WCL↗
MariBenchSoon
LLM × Physical WorldSoon

RQ4Deployment

Benchmarks and systems that carry models into scientific, maritime, wireless, and physical settings.

Publications iconSelected Publications

SubspacePath Pruner project figure

SubspacePath Pruner: Inference-time Pruning via Probe-based Representation-Parameter CouplingRQ1RQ2RQ4

Zhiren Gong, Yikun Hou, Fan Wu, Che Wang, Fuyao Zhang, Tiantong Wu, Yurong Hao, Jiaming Zhang, Yiyang Duan, Tiantong Wang, Fei Huang, Chau Yuen, Wei Yang Bryan Lim

SubspacePath Pruner is an inference-time, training-data-free structured pruner: a lightweight probe maps each scenario's representation subspace to a sparse, head-wise reasoning pathway, so the model keeps only the attention heads a given domain needs. This scenario-adaptive routing preserves accuracy while cutting computation, with no per-scenario fine-tuning.
# Model/Agent Efficiency # Inference-Time Pruning # Routing
CoCurve: overview of single-unit probing to H-conditioned joint pruning, and WikiText-2 perplexity across pruning ratios on six models

CoCurve: Cross-Module Co-Pruning Curvature for Structured LLM PruningRQ1RQ2RQ4

Zhiren Gong, Zeng Zihao, Tiantong Wang, Yixin Wang, Honoka Anada, Zijie Wang, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim

Most structured pruners score each unit in isolation, the wrong granularity for a Transformer's shared residual stream. CoCurve prunes attention and FFN units jointly via a Fisher matrix whose off-diagonal entries are cross-module co-pruning curvature edges, recoverable from just M single-unit ablations. It leads every capability class on Llama-3.1-8B at 20% sparsity and best preserves fragile code.
# Cross-Module Co-Pruning # Curvature Edges # Bridge Units
Model Consolidation under Budget: 2x2 results — accuracy across ViT scales, data-budget crossover, update-budget frontier, and distillation gain vs expert count

Model Consolidation under Budget: From Closed-Form Merging to End-to-End DistillationRQ2RQ4

Zhiren Gong, Tiantong Wang, Zeng Zihao, Yiyang Duan, Ming Xiao, Chau Yuen

Coming soon
Data-assisted model merging spans two regimes — cheap closed-form merging and compute-heavy end-to-end distillation (E2E-Merge) — both optimizing the same task-conditioned objective. We map when each is worth its cost: closed-form wins under tight data, compute, or memory; distillation pays once resources are ample, reaching 94.3% on ViT-L/14. Not a universal winner, but a practical accuracy–efficiency frontier.
# Model Merging # Budget–Accuracy Frontier # Distillation vs Closed-Form
SkillFocus: capability-space skill evolution vs existing execution-evidence skill evolution

SkillFocus: Evolving Agent Skills via Capability DecompositionRQ2RQ3RQ4

Ning Wang, Zhiren Gong, Bingdong Li, Peng Yang, Aimin Zhou

Preprint
Agent skill evolution revises reusable procedural guidance from execution evidence alone, leaving which recurring task requirements drive a revision implicit. SkillFocus decomposes task requirements into a capability space that stays fixed as the skill evolves, maps outcomes to the capability that leaves the most tasks unresolved, and targets revision there. Across four heterogeneous benchmarks it reaches the best held-out accuracy on all four — +5.7 points over the strongest competitor — using 24% fewer evolution tokens.
# Agent Skills # Capability Decomposition # Skill Evolution
SpecRoute overview: closed-loop multi-step speculation, prefix verification with KV reuse, and shared-GPU scheduling

SpecRoute: Multi-Step Speculation for Efficient LLM Agent ServingRQ2RQ4

Zeng Zihao, Siyi Li, Zhiren Gong, Lei Xiao, Wei Yang Bryan Lim

Coming soon
LLM agents interleave model reasoning with tool execution, leaving the multi-turn chain on the critical path. SpecRoute is a serving runtime that accelerates agents through multi-step speculative execution: a lightweight draft agent speculates on future tool actions and runs them off-GPU while the target processes its context, then verifies observations at tool-call boundaries with KV reuse, under a slack-aware scheduler that shares one GPU. It reaches 1.28–1.56× single-request speedups over vLLM with comparable task quality, and cuts mean completion time by 28–36% at low load under online serving.
# LLM Agent Serving # Speculative Execution # Inference Efficiency
SoT project figure

State of Thought Enables Endogenous ReasoningRQ1RQ3

Zhiren Gong, Yikun Hou, Zeng Zihao, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim

NeurIPS 2026
State of Thought (SoT) lets an LLM steer its own reasoning: a compact dynamics-geometric state read from the model's internal information transfer drives a 582-parameter controller that selects the sparse historical evidence each step needs and decides when to stop. Across 3 LLMs and a VLM it raises accuracy while cutting tokens 65.5% and latency 47.6%.
# Reasoning # Efficiency # State Modeling
Rule-Guided Active Inference figure

Learning Human Habits with Rule-Guided Active InferenceRQ3RQ4

Zhiren Gong, Chao Yang, Wendi Ren, Shuang Li

ICLR 2026
Learning Human Habits combines active inference with explicit rule constraints to model how habits form and drive behavior in sequential decision settings. The rule-guided formulation makes planning both more accurate and lower-latency in realistic environments, while keeping the learned policy interpretable.
# Active Inference # Decision Intelligence # Interpretable Planning
LoTR project figure

LoTR: Logic-of-Thought Routing for Plug-and-Play Reasoning of LLMsRQ1RQ3

Zhiren Gong, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim

Coming soon
External reasoning paradigms (CoT, Plan-and-Solve, Self-Refine, sampling, search) prescribe how to reason, but leave the model's internal computation pathway fixed. LoTR (Logic-of-Thought Routing) is a plug-and-play, inference-time module that, with frozen weights, reads a small basis of recurring logic states with a lightweight probe and mixes cached per-state attention-head templates to route each step through the pathway its logic needs. On Llama-3.1-8B (8 benchmarks × 8 paradigms) it lifts the macro by +2.66 pts (+5.37% relative) at only +0.15% tokens, with positive transfer to Qwen2.5-14B and Mixtral-8×7B.
# Plug-and-Play # Logic Routing # LLM Reasoning
CrossAgentMem overview figure

Concept-Profile Guided Cross-Agent Memory Sharing for Multi-Turn ConversationRQ2RQ3

Zijie Wang, Zhiren Gong, Zeng Zihao, Tiantong Wu, Lei Xiao, Wei Yang Bryan Lim

Coming soon
CrossAgentMem shares memory across agents in multi-turn conversation, guided by concept profiles that decide what is worth passing on. This memory-aware collaboration lifts multi-agent reasoning accuracy under a bounded token budget, controlling context overhead rather than letting it grow.
# Multi-Agent Reasoning # Context Efficiency # Memory Systems
ParetoEvo: Pareto-guided evolutionary tree search framework and three-objective results on TSP

ParetoEvo: Pareto-Guided Evolutionary Tree Search for Automatic Heuristic Design with LLM AgentsRQ2RQ3RQ4

Haoyun Li, Ming Xiao, Yong Liang Guan, Zhiren Gong

Coming soon
ParetoEvo casts automatic heuristic design as Pareto-guided evolutionary tree search: an LLM Actor proposes heuristics and a Reflector critiques them, while a Pareto-UCT search jointly optimizes solution quality, runtime efficiency, and worst-instance robustness. Across constructive and ACO-based routing and packing tasks (TSP, CVRP, MKP, BPP), no LLM baseline dominates it — it reaches higher hypervolume and lower IGD far earlier in the search.
# LLM Agents # Heuristic Design # Evolutionary Search
Conditional Co-Ablation project figure

Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer CircuitsRQ1

Zhiren Gong, He Lu, Tiantong Wang, Yichi Zhang, Yixin Wang, Zeng Zihao, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim

In transformer circuits, self-repair means the very ablation used to test a circuit can shift which components carry the behavior, so a circuit that explains the intact model may be incomplete. CoAx formalizes this as conditional circuit completion, ranking components by how much their effect grows once a primary set is removed — label-free and forward-only. It recovers GPT-2-small's IOI backups at 0.941 ROC-AUC and cuts circuit incompleteness from 0.75 to 0.21.
# Mechanistic Interpretability # Self-Repair # Circuit Completion
RLTB: a unified target-directed post-training framework — behavioral targets, output credit, and TB-GRPO

Reinforcement Learning toward Target BehaviorRQ1RQ4

Yichi Zhang, Zhiren Gong, Yucheng Chen, Jun Liu, Si Yong Yeo

Coming soon
RLTB reframes post-training as reaching an explicitly declared behavioral destination. Behavioral requirements — on individual outputs, related outputs, or responses to interventions — become point, bound, or interval targets that jointly define a target region; each behavior's deviation sets the direction and strength of its corrective feedback, which weakens near the target and vanishes upon satisfaction. TB-GRPO turns this into group-relative advantages with leave-one-prompt-group-out cross-fitting, reaching the smallest joint target gaps among baselines without case-specific retuning.
# Reinforcement Learning # Post-Training # Target Behavior
From Local Prediction Error to Final Outputs: local vs terminal geometry and trained-Transformer gains

From Local Prediction Error to Final Outputs: A Theory of Expert Sharing in Multi-Step ReasoningRQ1RQ3

Tiantong Wang, Zhiren Gong, Yixin Wang, Tiantong Wu, Wei Yang Bryan Lim

Coming soon
Sharing one expert across recurring reasoning steps can lower local prediction error yet raise final-output error. This work derives exact criteria for when local sharing gains survive composition — trading variance savings against sharing bias — characterizes a local-to-final risk reversal, and shows that predictive success alone cannot establish that an expert implements the operation attributed to it.
# Expert Sharing # Multi-Step Reasoning # Theory
Paper4 visual

Why Mixture of Experts Needs Fewer Samples Than Dense Networks: The Information Exponent Collapse TheoremRQ1RQ2

Tiantong Wang, Zhiren Gong, Yikun Hou, Tiantong Wu, Wei Yang Bryan Lim

Coming soon
This work develops a theory of sample efficiency for Mixture-of-Experts: geometry-aware routing lets sparse MoE models cross learning barriers with far fewer samples than dense networks, through an information-exponent collapse that lowers the effective complexity of the target function.
# Mixture of Experts # Sample Complexity # Theory
XDomainBench project figure

XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge CompositionRQ3RQ4

Zhiren Gong, Tiantong Wu, Jiaming Zhang, Fuyao Zhang, Che Wang, Yurong Hao, Yikun Hou, Foo Ping, Yilei Zhao, Fei Huang, Chau Yuen, Wei Yang Bryan Lim

XDomainBench stress-tests large models on cross-domain scientific knowledge composition — combining concepts across fields. It systematically reveals where and why reasoning collapses as compositional complexity grows, turning a diffuse failure mode into a measurable, diagnosable benchmark.
# Benchmark # Reliability # Scientific Reasoning
MariBench evaluation protocol: operator question, LLM agent calling AIS tools over hidden records, structured output, and outcome plus process scoring

MariBench: Benchmarking LLM Agents for Maritime Operational Judgment over Raw AIS DataRQ3RQ4

He Lu, Zhiren Gong, Fuyao Zhang, Tianbing Xia, Yufan Huang, Wei Yang Bryan Lim, Ran Yan

Coming soon
MariBench is the first interactive agent benchmark for maritime operational judgment over raw AIS — 380 real cases hidden behind callable tools, scored on both process and answer across three levels. Across 17 models, grounding is near-ceiling (79–100%) yet judgment stalls at 45–62%: the dominant failure is over-escalation, flagging benign artifacts as anomalies — which scale doesn't fix but human experts avoid.
# Agent Benchmark # Operational Judgment # Over-Escalation
Paper1 visual

Channel Gains to Captions: Task-Unified Multi-Level RF Sensing with Vision-Language ModelsRQ4

Tianyu Hu, Zhiren Gong, Haowei Cui, Shuai Wang, Samson Lasaulce, Lingxiang Li, Wassim Hamidouche, Zhi Chen, Merouane Debbah

IEEE Wireless Communications Letters
Channel Gains to Captions unifies multi-level RF sensing with vision-language models, aligning channel-level signal structure with caption-level semantics in a single task-unified framework — one model that reads wireless signals and describes what they mean.
# RF Sensing # Vision-Language Models # Multimodal Intelligence
Paper2 visual

3D Radio Map Reconstruction based on Generative Adversarial Networks under Constrained Aircraft TrajectoriesRQ4

Tianyu Hu, Yang Huang, Junting Chen, Qihui Wu, Zhiren Gong

IEEE TVT 2023
A GAN-based method reconstructs dense 3D radio maps from sparse aircraft measurements collected under constrained flight trajectories, recovering signal-strength fields that limited sampling misses and improving reconstruction quality under realistic sensing conditions.
# Wireless AI # 3D Radio Map # GAN
C2U in RAG visual

C2U in RAG: Compositional Concept Unlearning in Retrieval-Augmented GenerationRQ4

Yiyang Duan, Fan Wu, Tiantong Wu, Che Wang, Zhiren Gong, Wei Guo, Yang Cao, Wei Yang Bryan Lim

Coming soon
C2U in RAG studies compositional concept unlearning for retrieval-augmented generation: it selectively removes targeted knowledge — including concepts that emerge only from combinations — while preserving utility on everything meant to be retained.
# RAG # Unlearning # Compositionality
Paper3 visual

When Large Language Models Meet the Physical World: Beyond Task-Level, Toward Situation-LevelRQ4

Yang Liu, Zhiren Gong, Xiaoping Wang, Jianbo Zheng, Guanghui Ye, Jinhong Hu, Kai Lu, Wei Yang Bryan Lim

Coming soon
When LLM systems meet the physical world, task-level competence breaks down on non-semantic, embodied signals. This work argues for moving beyond task-level toward situation-level understanding — modeling the broader situation a decision sits in — to build more robust and reliable physical-world cognition.
# LLM + Physical World # Robustness # Reliability

Achievement iconAchievement

  • NTU Research Scholarship, PhD Program.
  • NUAA Outstanding Graduate (Top 1%) and Outstanding Thesis 2023 (Top 4%).
  • Headmaster's Special Mention (2020, 2022), highest undergraduate honor at NUAA.
  • Top 10 Outstanding Youth of NUAA (2022).
  • Aviation Industry Honors Scholarship (Top 0.01%).
  • National Scholarship of China (Top 1%).
  • First Prize Outstanding Student Scholarship (Top 5%, 2019 - 2023).
  • National Second Prize (Top 3%), China Undergraduate Mathematical Contest in Modeling (CUMCM), 2021.
  • Meritorious Winner (Top 8%), Mathematical Contest in Modeling (MCM), 2021.
  • Asia Second Prize, Asia Pacific Mathematical Contest in Modeling (APMCM), 2020 and 2021.
  • Provincial Second Prize, Jiangsu Province Engineering Ability Competition, 2021.
  • National First Prize (Top 10%), National College Students Academic Physics Competition, 2020.
  • National Second Prize, East China Mathematical Contest in Modeling, 2020.
  • National Second Prize, Huajiao Cup National College Students Mathematics Competition, 2020.
  • National Second Prize (Top 20%), National College Students Mathematics Competition, 2019.
  • Three-dimensional Spectrum Situation Completion Method and Apparatus based on Generative Adversarial Network
    PCT/CN2022/073723, WO 2022/206149 A1 (co-inventor).
  • An Unmanned Aerial Vehicle Decision Evaluation Method in a Complex Environment
    Chinese National Invention Patent CN113673149 (A).
  • Adaptive Parking Lot Intelligent Management and Control Method based on Neural Network
    Chinese National Invention Patent CN113283302 (A).
  • Intelligent Garbage Sorting Device and Method based on Deep Learning Technology
    Chinese National Invention Patent CN113562355 (A).
  • Liquid Surface Tension Meter based on Intelligent Control of an STM32 Single-chip Microcomputer, and Use Method thereof
    Chinese National Invention Patent CN113376059 (A).
  • Device and Method for Achieving Energy-saving Cooling of Automobile by Utilizing Air Outlet Frame
    Chinese National Invention Patent CN112428788 (A).

Service iconService

Reviewer · Conference: ICML 2026 (Gold Reviewer), NeurIPS 2026, AAAI 2027, ICLR 2027, ICIC 2026, ICML AdaptFM Workshop.

Reviewer · Journal: Transactions on Machine Learning Research (TMLR).

Teaching Assistant: Data Science Foundations (SC3021), NTU, 2026 Spring; Excellent Support (Top 10% of Lab TAs).