Research
* Equal contribution. † Project lead. ‡ Corresponding author.
2026
-
MARCH: Scaling Recurrent Memory with Content-Routed State Anchors
-
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
-
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
-
Qwen-AgentWorld: Language World Models for General Agents
-
ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics
-
Draft-OPD: On-Policy Distillation for Speculative Draft Models
-
Post-Trained MoE Can Skip Half Experts via Self-Distillation
-
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling
-
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning
-
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
-
How Far Can Unsupervised RLVR Scale LLM Training?
-
P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads
2025
-
Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation
-
JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
-
Accurate de novo sequencing of the modified proteome with OmniNovo
-
P1: Mastering Physics Olympiads with Reinforcement Learning
-
DePass: Unified Feature Attributing by Simple Decomposed Forward Pass
-
V-GameGym: Visual Game Generation for Code Large Language Models
-
FlowRL: Matching Reward Distributions for LLM Reasoning
-
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
-
A Survey of Reinforcement Learning for Large Reasoning Models
-
Towards a Unified View of Large Language Model Post-Training
-
From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery
-
SSRL: Self-Search Reinforcement Learning
-
Automating Exploratory Multiomics Research via Language Models
-
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
-
Self-Reflective Reinforcement Learning for Diffusion-based Image Reasoning Generation
-
When super-resolution meets camouflaged object detection: A comparison study.
-
Fourier Position Embedding: Enhancing Attentions Periodic Extension for Length Generalization
-
Free Process Rewards without Process Labels
-
How to Synthesize Text Data without Model Collapse?
-
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
-
TTRL: Test-Time Reinforcement Learning
-
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
-
A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond.
-
Technologies on Effectiveness and Efficiency: A Survey of State Spaces Models
-
Less is More: Efficient Model Merging with Binary Task Switch
-
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
-
Process reinforcement through implicit rewards.
-
OpenPRM: Building Open-domain Process-based Reward Models with Preference Trees
-
Advancing LLM Reasoning Generalists with Preference Trees
2024
-
Towards AI-45° Law: A Roadmap to Trustworthy AGI
-
Automating exploratory proteomics research via language models.
-
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines.
-
Empowering private tutoring by chaining large language models.
-
Neural Residual Diffusion Models for Deep Scalable Vision Generation
-
Efficient Diffusion Models: A Comprehensive Survey from Principles to Practices
-
LFQA-E: Carefully Benchmarking Long-form QA Evaluation
-
MSI-Agent: Incorporating Multi-Scale Insight into Embodied Agents for Superior Planning and Decision-Making
-
On the token distance modeling ability of higher RoPE attention dimension
-
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
-
EVA-Score: Evaluating Abstractive Long-form Summarization on Informativeness through Extraction and Validation
-
Towards Building Specialized Generalist AI with System 1 and System 2 Fusion
-
Safe-SD: Safe and Traceable Stable Diffusion with Text Prompt Trigger for Invisible Generative Watermarking
-
Exploring Adversarial Robustness of Deep State Space Models
-
UltraMedical: Building Specialized Generalists in Biomedicine
-
Enhancing Adversarial Transferability via Information Bottleneck Constraints
-
Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
-
Intuitive Fine-Tuning: Towards Simplifying Alignment into a Single Process
-
Fast and Slow Generating: An Empirical Study on Large and Small Language Models Collaborative Decoding
-
SMR: State Memory Replay for Long Sequence Modeling
-
CoGenesis: A Framework Collaborating Large and Small Language Models for Secure Context-Aware Instruction Following
-
Trust in Internal or External Knowledge? Generative Multi-Modal Entity Linking with Knowledge Retriever
-
On Large Language Models' Hallucination with Regard to Known Facts
-
PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning
-
Contrastive Augmented Graph2Graph Memory Interaction for Few Shot Continual Learning
-
Investigating Deep Watermark Security: An Adversarial Transferability Perspective
-
Large Language Models as Biomedical Hypothesis Generators: A Comprehensive Evaluation
-
Interactive Continual Learning: Fast and Slow Thinking
-
LAKE-RED: Camouflaged Images Generation by Latent Background Knowledge Retrieval-Augmented Diffusion
-
Generative AI for Complex Scenarios: Language Models are Sequence Processors
-
Generative Multi-Modal Knowledge Retrieval with Large Language Models
-
LMD: Faster Image Reconstruction with Latent Masking Diffusion
-
AdapEdit: Spatio-Temporal Guided Adaptive Editing Algorithm for Text-Based Continuity-Sensitive Image Editing
2023
-
Large Language Models are Zero Shot Hypothesis Proposers
-
CRaSh: Clustering, Removing, and Sharing Enhance Fine-tuning without Full Large Language Model
-
Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
-
Sparse Low-rank Adaptation of Pre-trained Language Models
-
Improving Robustness of Intent Detection Under Adversarial Attacks: A Geometric Constraint Perspective
-
SAM Struggles in Concealed Scenes -- Empirical Study on Segment Anything
-
Trustworthy AI: From Principles to Practices