Publications Zhiren Gong

Home
Tags mark the research questions each paper tackles:RQ1 internal structureRQ2 model / agent efficiencyRQ3 reasoning efficiencyRQ4 deployment

Reasoning Efficiency

Making the reasoning process itself cheaper and more adaptive.

Methods that let a frozen model steer its own reasoning, selecting the right evidence, routing the right internal pathway, and deciding when to stop, so quality rises while tokens and latency fall.

Published & Preprints 2

State of Thought Enables Endogenous ReasoningRQ1RQ3

Zhiren Gong, Yikun Hou, Zeng Zihao, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim

NeurIPS 2026

A 582-parameter controller reads the model's own internal state to pick evidence and stop early, raising accuracy while cutting roughly 62% of generated tokens and 45% of latency across 3 LLMs and 2 VLMs.

Learning Human Habits with Rule-Guided Active InferenceRQ3RQ4

Zhiren Gong, Chao Yang, Wendi Ren, Shuang Li

ICLR 2026

Wake-and-sleep active inference compiles recurring situations into habit rules, so familiar contexts act instantly and novel ones fall back on full planning, reaching 91.3% Acc@3 at 36 ms versus 53 ms.

Coming soon 3

LoTR: Logic-of-Thought Routing for Plug-and-Play Reasoning of LLMsRQ1RQ3

Zhiren Gong, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim

Coming soon

Plug-and-play logic-conditioned routing reads the step's logic state and re-blends cached head templates with frozen weights, adding 2.66 points on an 8x8 Llama study at only 0.15% more tokens.

Concept-Profile Guided Cross-Agent Memory Sharing for Multi-Turn ConversationRQ2RQ3

Zijie Wang, Zhiren Gong, Zeng Zihao, Tiantong Wu, Lei Xiao, Wei Yang Bryan Lim

Coming soon

Concept-profile-guided memory sharing across agents lifts multi-agent reasoning accuracy under a bounded token budget, controlling context growth instead of letting it balloon.

ParetoEvo: Pareto-Guided Evolutionary Tree Search for Automatic Heuristic Design with LLM AgentsRQ2RQ3RQ4

Haoyun Li, Ming Xiao, Yong Liang Guan, Zhiren Gong

Coming soon

Casts automatic heuristic design as Pareto-guided evolutionary tree search, where an LLM actor and reflector jointly optimize solution quality, runtime, and worst-instance robustness.

Model / Agent Efficiency

Shrinking and serving the model before it ever reasons.

Inference-time pruning, cross-module co-pruning, budget-aware merging, speculative agent serving, and reusable agent skills, bringing efficiency at the parameter and system level without retraining the backbone.

Published & Preprints 3

SubspacePath Pruner: Inference-time Pruning via Probe-based Representation-Parameter CouplingRQ1RQ2RQ4

Zhiren Gong, Yikun Hou, Fan Wu, Che Wang, Fuyao Zhang, Tiantong Wu, Yurong Hao, Jiaming Zhang, Yiyang Duan, Tiantong Wang, Fei Huang, Chau Yuen, Wei Yang Bryan Lim

ICML 2026

A lightweight probe maps each scenario's domain subspace to a sparse head pathway, so a frozen model keeps only the heads a domain needs and prunes at inference time with no fine-tuning or scenario data.

CoCurve: Cross-Module Co-Pruning Curvature for Structured LLM PruningRQ1RQ2RQ4

Zhiren Gong, Zeng Zihao, Tiantong Wang, Yixin Wang, Honoka Anada, Zijie Wang, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim

Preprint

Prunes attention and FFN jointly through cross-module co-pruning curvature recovered from just M single-unit probes, leading every capability class on Llama-3.1-8B at 20% sparsity.

SkillFocus: Evolving Agent Skills via Capability DecompositionRQ2RQ3RQ4

Ning Wang, Zhiren Gong, Bingdong Li, Peng Yang, Aimin Zhou

Preprint

Targets skill revision at the capability that leaves the most tasks unsolved, reaching the best held-out accuracy on all four benchmarks, 5.7 points over the strongest baseline, with 24% fewer tokens.

Coming soon 2

Model Consolidation under Budget: From Closed-Form Merging to End-to-End DistillationRQ2RQ4

Zhiren Gong, Tiantong Wang, Zeng Zihao, Yiyang Duan, Ming Xiao, Chau Yuen

Coming soon

Charts when cheap closed-form merging beats compute-heavy end-to-end distillation, mapping a clear accuracy and efficiency frontier that reaches 94.3% on ViT-L/14.

SpecRoute: Multi-Step Speculation for Efficient LLM Agent ServingRQ2RQ4

Zeng Zihao, Siyi Li, Zhiren Gong, Lei Xiao, Wei Yang Bryan Lim

Coming soon

A draft agent speculates future tool actions off-GPU and verifies them at tool-call boundaries with KV reuse, giving 1.28 to 1.56x single-request speedups over vLLM at matched task quality.

Theory & Interpretability

Understanding what the computation inside a model actually does.

Mechanistic accounts of transformer circuits, expert sharing, and mixture-of-experts sample efficiency, turning internal structure into claims we can test rather than trust.

Published & Preprints 1

Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer CircuitsRQ1

Zhiren Gong, He Lu, Tiantong Wang, Yichi Zhang, Yixin Wang, Zeng Zihao, Ming Xiao, Chau Yuen, Wei Yang Bryan Lim

Preprint

Ranks components by how much their ablation effect grows once a primary set is removed, recovering GPT-2 IOI self-repair backups at 0.941 ROC-AUC and cutting circuit incompleteness from 0.75 to 0.21.

Coming soon 3

Reinforcement Learning toward Target BehaviorRQ1RQ4

Yichi Zhang, Zhiren Gong, Yucheng Chen, Jun Liu, Si Yong Yeo

Coming soon

Reframes post-training as reaching a declared behavioral target region, where TB-GRPO turns each behavior's deviation into group-relative advantages and reaches the smallest joint target gaps among baselines.

From Local Prediction Error to Final Outputs: A Theory of Expert Sharing in Multi-Step ReasoningRQ1RQ3

Tiantong Wang, Zhiren Gong, Yixin Wang, Tiantong Wu, Wei Yang Bryan Lim

Coming soon

Proves when sharing one expert across reasoning steps lowers local error yet raises final-output error, characterizing a local-to-final risk reversal that predictive success alone cannot rule out.

Why Mixture of Experts Needs Fewer Samples Than Dense Networks: The Information Exponent Collapse TheoremRQ1RQ2

Tiantong Wang, Zhiren Gong, Yikun Hou, Tiantong Wu, Wei Yang Bryan Lim

Coming soon

Proves mixture-of-experts crosses learning barriers with far fewer samples than dense networks, through an information-exponent collapse that lowers the target function's effective complexity.

Applications

Putting foundation models to work, and stress-testing where they break.

Benchmarks and systems that take models into scientific composition, maritime judgment, RF sensing, and the physical world, exposing the failure modes real deployment must handle.

Published & Preprints 3

XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge CompositionRQ3RQ4

Zhiren Gong, Tiantong Wu, Jiaming Zhang, Fuyao Zhang, Che Wang, Yurong Hao, Yikun Hou, Foo Ping, Yilei Zhao, Fei Huang, Chau Yuen, Wei Yang Bryan Lim

ICML 2026

Stress-tests cross-domain scientific knowledge composition and shows exactly where and why reasoning collapses as the number of fused domains grows, turning a diffuse failure into a measurable axis.

Channel Gains to Captions: Task-Unified Multi-Level RF Sensing with Vision-Language ModelsRQ4

Tianyu Hu, Zhiren Gong, Haowei Cui, Shuai Wang, Samson Lasaulce, Lingxiang Li, Wassim Hamidouche, Zhi Chen, Merouane Debbah

IEEE WCL

Unifies multi-level RF sensing with vision-language models so one model reads channel-level wireless structure and describes it in caption-level semantics within a single task-unified framework.

3D Radio Map Reconstruction based on Generative Adversarial Networks under Constrained Aircraft TrajectoriesRQ4

Tianyu Hu, Yang Huang, Junting Chen, Qihui Wu, Zhiren Gong

IEEE TVT 2023

A GAN reconstructs dense 3D radio maps from sparse aircraft measurements collected under constrained flight trajectories, recovering signal fields that limited sampling misses.

Coming soon 3

MariBench: Benchmarking LLM Agents for Maritime Operational Judgment over Raw AIS DataRQ3RQ4

He Lu, Zhiren Gong, Fuyao Zhang, Tianbing Xia, Yufan Huang, Wei Yang Bryan Lim, Ran Yan

Coming soon

The first interactive agent benchmark for maritime judgment over raw AIS: across 17 models grounding is near-ceiling at 79 to 100%, yet judgment stalls at 45 to 62%, dominated by over-escalation.

C2U in RAG: Compositional Concept Unlearning in Retrieval-Augmented GenerationRQ4

Yiyang Duan, Fan Wu, Tiantong Wu, Che Wang, Zhiren Gong, Wei Guo, Yang Cao, Wei Yang Bryan Lim

Coming soon

Compositional concept unlearning for retrieval-augmented generation removes targeted and even emergent combined knowledge while preserving utility on everything meant to be retained.

When Large Language Models Meet the Physical World: Beyond Task-Level, Toward Situation-LevelRQ4

Yang Liu, Zhiren Gong, Xiaoping Wang, Jianbo Zheng, Guanghui Ye, Jinhong Hu, Kai Lu, Wei Yang Bryan Lim

Coming soon

Argues that task-level competence breaks on embodied, non-semantic signals and calls for situation-level understanding to build robust, reliable physical-world cognition.