alphaXiv

Explore

Researchers

Sign In

MCP Server

Autoresearch

Browser Extension

BlogSend Feedback?

Follow the latest research

alphaXiv connects papers, researchers, and organizations, grounding its answers in the underlying work.

Alt + Enter to search
Sign up

Finite Time Blowup for Navier–Stokes

OpenAI

A smooth, compactly forced three-dimensional flow can develop unbounded velocity in finite time while retaining uniformly bounded kinetic energy.

08 Sept 2026
5kviews
On the Navier–Stokes Millennium Prize Problem
OpenAI

A coordinated multiagent search produced a proposed resolution and Lean formalization, giving researchers artifacts to inspect, verify, and build upon.

08 Sept 2026
1kviews
On the Navier–Stokes Millennium Prize Problem

SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators

UMass AmherstUC Berkeley
Yuncong YangZhengtao HanYilun DuYilun Du

A brief visual calibration lets robot world models simulate actions in unseen camera views and embodiments without parameter updates or additional training.

08 Sept 2026

Researchers to follow

View all
Ian Goodfellow

Ian Goodfellow

Co-Founder @ Stealth Startup, Previously Research Scientist @ DeepMind

Yoshua Bengio

Yoshua Bengio

President and Scientific Director @ LawZero, Founder and Scientific Advisor @ Mila - Quebec Artificial Intelligence Institute, Canada CIFAR AI Chair @ CIFAR, Full Professor, CS @ Université de Montréal

Andrew Ng

Andrew Ng

Managing Partner @ AI Aspire, Managing General Partner @ AI Fund, Founder @ DeepLearning.AI, Adjunct Professor, CS @ Stanford University, Chairman and Co-Founder @ Coursera

Kaiming He

Kaiming He

Distinguished Scientist @ Google DeepMind, Associate Professor, EECS @ MIT

Sergey Levine

Sergey Levine

Co-Founder @ Physical Intelligence, Associate Professor, EECS @ UC Berkeley

Demis Hassabis

Demis Hassabis

Chair @ Google DeepMind, Chief Scientist @ Alphabet, Founder & CEO @ Isomorphic Labs

Ilya Sutskever

Ilya Sutskever

CEO and Co-Founder @ Safe Superintelligence Inc, Previously Co-Founder and Chief Scientist @ OpenAI

Geoffrey Hinton

Geoffrey Hinton

Emeritus Professor, CS @ University of Toronto, Previously Vice President and Engineering Fellow @ Google

Are you a researcher? Find your profile

AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing

SJTUShanghai Innovation Institute
Ziyang MaZiyang MaZhikang NiuXie ChenXie Chen

A single open-source model can generate, clone, edit, enhance, and separate speech through natural-language instructions and optional audio context.

08 Sept 2026
235views

OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining

NUSTsinghua
Yuran WangSiqiao HuangShanghang ZhangShanghang Zhang

A modular research stack enables controlled comparisons of world–action architectures, training recipes, and data mixtures across simulations and physical robots.

07 Sept 2026
463views

Omni Interaction Agent Technical Report

TencentZJU
OrantqingShengpeng JiXiaoyu ShenXiaoyu Shen

In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander continuously receives streaming inputs across multiple modalities, including video, speech, and text, enabling natural full-duplex interaction in both everyday conversations and complex workflow-oriented agent scenarios. Users can interrupt the model at any time, while the model can also proactively provide intermediate feedback or ask follow up questions. To natively support these capabilities, Gander adopts two key architectural designs: 1) It employs a Cerebellum-Brain collaborative framework, in which the Cerebellum is responsible for realtime interaction and omni conversational capabilities, while the Brain handles complex reasoning and higher-level agentic tasks. The two components interact continuously through tool calling and the agent orchestration runtime. 2) The Cerebellum is built upon a streaming Thinker-Talker architecture, user inputs and model outputs are further flattened into an ordered token stream at the chunk level, providing a unified representation for low latency, continuous interaction. We conduct comprehensive evaluations of Gander across four dimensions: conversational ability, omni understanding, interactive capability, and agentic intelligence. Internal human evaluations demonstrate that Gander maintains the natural and expressive spoken dialogue capabilities of SOTA open source models while achieving competitive performance in omni interaction. Gander also demonstrates robustness in challenging real-world scenarios, including background noise interference, multi-party interactions, and backchannel communication. We release Gander together with its models, code, and data to facilitate further research and development in the community.

09 Sept 2026
168views

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

KAISTMicrosoft
Youngrok ParkSangmin BaeSangmin BaeAaron CourvilleAaron Courville

Weak teachers can accelerate stronger models’ verifier-based training while letting them surpass the teachers and retain their own reasoning strategies.

08 Sept 2026

Procedural Graphs: Self-Evolving Execution Structures for LLM Agents

Google
Yuxing LuYicheng ChenShanchan Wu

Editable procedural graphs help language-model agents choose valid next steps, avoid repetitive tool loops, and improve execution plans from trajectory feedback.

08 Sept 2026
127views

IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications

Yale
Yiling MaYilun ZhaoYilun ZhaoArman CohanArman Cohan

A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. We study the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. We construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts. We introduce IdeaAMBIG, a benchmark of 660 evidence-grounded instances: 163 real-world gaps from reproducibility reports and GitHub issues, and 497 controlled synthetic gaps injected into codification-ready references. IdeaAMBIG evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Defect localization receives only the specification, whereas clarification additionally receives the annotated defect. Across 13 LLMs, the best model achieves 9.6% Macro Defect Recovery Rate on real-world instances but 80.6% Macro Clarification Action Success Rate when given the defect. In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%. Across all evaluated models, defect localization is the main bottleneck, with stronger clarification given the defect.

09 Sept 2026

Point4D: Long-range 4D Motion Reconstruction

CMU
Minsik JeonJay KarhadeDeva RamananDeva Ramanan

3D queries let feed-forward models maintain dense 3D point trajectories across hundreds of video frames, even through occlusion and field-of-view changes.

08 Sept 2026

Long-Horizon Language Model Reinforcement Learning via Progressive Point Matching

UC BerkeleyCMU
PF
Preston Fu
KF
Kevin Frans
Sergey LevineSergey Levine

Segment-level rewards from a single reference trajectory let language models learn from partial progress on difficult reasoning tasks where outcome rewards provide no signal.

07 Sept 2026

Researchers to follow

View all
Linxi "Jim" Fan

Linxi "Jim" Fan

Director & Distinguished Research Scientist @ NVIDIA, Previously CS PhD Student @ Stanford University

Oriol Vinyals

Oriol Vinyals

Co-Founder @ Discovery Loop, Previously VP of Research @ Google DeepMind

Terence Tao

Terence Tao

Distinguished Professor @ UCLA, Previously Mathematics PhD Student @ Princeton University

Tri Dao

Tri Dao

Assistant Professor, CS @ Princeton University, Co-Founder & Chief Scientist @ Together AI

Quoc V. Le

Quoc V. Le

Co-Founder @ Discovery Loop, Previously Research Scientist @ Google

Noam Shazeer

Noam Shazeer

Previously VP Engineering @ Google

Song Han

Song Han

Distinguished Scientist, Director of Efficient AI Research @ NVIDIA, Associate Professor, EECS @ Massachusetts Institute of Technology

Junyang Lin

Junyang Lin

Independent Researcher @ Unaffiliated, Previously Tech Lead @ Alibaba Group

RenderFormer-V2: Neural Rendering with Heterogeneous Scene Primitives

StanfordMicrosoft Research
Chong ZengChong ZengYue DongLvmin ZhangLvmin Zhang

A single pretrained renderer handles complex scenes with mixed geometry, materials, environment lighting, caustics, and volumetric scattering without per-scene training.

04 Sept 2026

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness

NeoHorse TeamGuoliang CaoGuohao Dai

Routing harnesses turn deployed agent trajectories and capability feedback into targeted training data, enabling models to improve on interactive tasks.

08 Sept 2026
182views

Miles v0.1: Production-Level Post-Training

RadixArkTom ChenMao Cheng

Researchers can train frontier-scale language and diffusion models with customizable synchronous or asynchronous post-training workflows across NVIDIA and AMD hardware.

08 Sept 2026
130views3k

Sparks of In Silico Cognitive Science: Theories from Simulated Data Can Generalize to Humans

PrincetonCornell
Akshay K. JagadishYounes StrittmatterThomas L. GriffithsThomas L. Griffiths

An automated theory-discovery loop using simulated behavior produced decision-making theories that predicted held-out human choices better than canonical models.

07 Sept 2026

Rethinking Safety for Generalist Robots

StanfordUCLA
Rohan SinhaAnushri DixitAnirudha MajumdarAnirudha Majumdar

The taxonomy broadens robot safety beyond collisions to include semantic hazards, misalignment, misuse, privacy, and risks spanning development through deployment.

06 Sept 2026

3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints

TsinghuaETHZ
Ziqin HuangYingyue LiMasayoshi TomizukaMasayoshi Tomizuka

Multi-view waypoint predictions triangulate into executable 3D trajectories, helping manipulation policies transfer to unfamiliar tasks, objects, and robot embodiments.

08 Sept 2026

A Brain-inspired Hierarchical Framework for Zero-Shot Robot Task Reasoning and Execution

CambridgeMIT
Guangming WangPengfei YeHaonan ChenHaonan Chen

Robots can adapt long-horizon manipulation plans to changing object states, recover from execution errors, and verify progress without task-specific training.

05 Sept 2026

The Interface of Theseus: The Rise of Just-In-Time Interfaces

Stanford
Michael S. BernsteinMichael S. Bernstein

The paper envisions interfaces generated on demand from user goals and context, enabling bespoke interactions that connect information across traditionally separate applications.

06 Sept 2026

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

PKUBAAI
Yankai FuNing ChenShanghang ZhangShanghang Zhang

Tactile feedback and future contact prediction help dexterous robots perform fine-grained manipulation while generalizing across clutter, lighting changes, and unfamiliar objects.

08 Sept 2026
There are no more papers matching your filters at the moment.
Sign in

Assistant