---
title: "τ₀-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation"
authors:
  - Xiaowei Cai
  - Yunuo Cai
  - Bingao Chen
  - Jingxiao Chen
  - Zhi Chen
  - Siyuan Feng
  - Tengyu Hou
  - Jingshun Huang
  - Han Jiang
  - Runkun Ju
  - Dong Li
  - Mingxiang Li
  - Shaowei Li
  - Xinchen Li
  - Yifan Li
  - Yi Liu
  - Zhongyuan Liu
  - Jianlan Luo
  - Junwen Miao
  - Ruiqi Ni
  - Buqing Nie
  - Mingjie Pan
  - Xinlin Ren
  - Jianheng Song
  - Jiaxu Wang
  - Peiqi Wang
  - Sen Wang
  - Xiaoyan Wang
  - Dafeng Wei
  - Dongming Wu
  - Pengwei Xie
  - Pu Yang
  - Hangjian Ye
  - Xiangyu Yue
  - Jinyu Zhang
  - Qinglin Zhang
  - Xueyong Zhao
  - Yue Zhou
release_date: 2026-07-27
source_url: https://tau0-vla.github.io/tau0-vla.pdf
status: "Technical report; no arXiv or peer-reviewed venue record verified on 2026-07-28"
conversion: "PyMuPDF4LLM 1.28.0 from the official rendered PDF; equations and complex figures may be preserved as extracted PNG assets"
---

> Conversion note: this Markdown was extracted from the official 17-page PDF. The PDF remains the authoritative rendering; equations and complex figure regions that did not convert cleanly are preserved under `assets/`.

# τ₀-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

Xiaowei Cai<sup>2</sup> , Yunuo Cai<sup>1,2</sup> , Bingao Chen<sup>2,3</sup> , Jingxiao Chen<sup>2</sup> , Zhi Chen<sup>2</sup> , Siyuan Feng<sup>2</sup> , Tengyu Hou<sup>2</sup> , Jingshun Huang<sup>1,2</sup> , Han Jiang<sup>2</sup> , Runkun Ju<sup>2</sup> , Dong Li<sup>2</sup> , Mingxiang Li<sup>2</sup> , Shaowei Li<sup>2</sup> , Xinchen Li<sup>2</sup> , Yifan Li<sup>1,2</sup> , Yi Liu<sup>1,2</sup> , Zhongyuan Liu<sup>2</sup> , Jianlan Luo<sup>1,2</sup> , Junwen Miao<sup>2</sup> , Ruiqi Ni<sup>2</sup> , Buqing Nie<sup>2</sup> , Mingjie Pan<sup>1,2</sup> , Xinlin Ren<sup>2</sup> , Jianheng Song<sup>2</sup> , Jiaxu Wang<sup>2,3</sup> , Peiqi Wang<sup>2</sup> , Sen Wang<sup>2</sup> , Xiaoyan Wang<sup>2</sup> , Dafeng Wei<sup>2</sup> , Dongming Wu<sup>2,3</sup> , Pengwei Xie<sup>2</sup> , Pu Yang<sup>2</sup> , Hangjian Ye<sup>1,2</sup> , Xiangyu Yue<sup>2,3</sup> , Jinyu Zhang<sup>1,2</sup> , Qinglin Zhang<sup>2</sup> , Xueyong Zhao<sup>2</sup> , Yue Zhou<sup>2</sup>

> 1Shanghai Institute of Innovation 2Agibot Finch 3The Chinese University of Hong Kong https://tau0-vla.github.io/


![](assets/paper_tau0-vla-2026.pdf-0001-03.png)


<!-- Start of picture text -->
Pretraining Data<br>Teleop Data<br>Autonomous Data<br>UMI Data<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0001-04.png)


<!-- Start of picture text -->
✓<br>Pick Up Spoon<br>P W V<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0001-05.png)


<!-- Start of picture text -->
✗<br>Pick Up Cup<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0001-06.png)



![](assets/paper_tau0-vla-2026.pdf-0001-07.png)


<!-- Start of picture text -->
P W V<br>✗<br>Place Cup<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0001-08.png)


<!-- Start of picture text -->
✗<br>Place Cup<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0001-09.png)


<!-- Start of picture text -->
✗<br>Pick Up Cup<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0001-10.png)


<!-- Start of picture text -->
✓<br>Pick Up Cup<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0001-11.png)


<!-- Start of picture text -->
P W V<br>✓<br>Place Cup<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0001-12.png)


<!-- Start of picture text -->
✗<br>Pick Up Spoon<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0001-13.png)



![](assets/paper_tau0-vla-2026.pdf-0001-14.png)


<!-- Start of picture text -->
Long-Horizon<br>Mobile Manipulation<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0001-15.png)



![](assets/paper_tau0-vla-2026.pdf-0001-16.png)



![](assets/paper_tau0-vla-2026.pdf-0001-17.png)



![](assets/paper_tau0-vla-2026.pdf-0001-18.png)



![](assets/paper_tau0-vla-2026.pdf-0001-19.png)



![](assets/paper_tau0-vla-2026.pdf-0001-20.png)



![](assets/paper_tau0-vla-2026.pdf-0001-21.png)



![](assets/paper_tau0-vla-2026.pdf-0001-22.png)


Fig. 1: **Overview of** τ₀ **-VLA.** The high-level policy uses world-model-guided test-time computation to search over subtask sequences. At each expansion step, a VLM proposes candidate subtasks, a world model predicts their visual outcomes, and a value model evaluates the resulting task progress. Beam search retains promising branches, after which a reflective model commits to the next subtask based on the retained candidates and their predicted consequences. The figure illustrates a single candidate expansion with _N_ = 3; recursive beam expansion to depth _D_ is omitted for brevity. The selected subtask then conditions a low-level VLA policy trained on 40,115 hours of heterogeneous real-world data, enabling long-horizon mobile manipulation, precise execution, and deployment across multiple robot embodiments.

**_Abstract_ —Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-languageaction (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce** τ₀ **VLA, a hierarchical robot foundation model that formulates highlevel subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives**

**before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distributionshifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.**

## I. INTRODUCTION

Long-horizon manipulation encompasses a broad class of real-world robotic tasks, such as cleaning a room, cooking a meal, or preparing a drink. These tasks require coherent

<!-- End of PDF page 1 -->

sequences of subtasks carried out over minutes or hours. The robot must locate and manipulate objects, interact with articulated structures, verify intermediate outcomes, and recover from local failures. Such tasks are not simply long sequences of motor commands; they are sequences of consequential decisions. Throughout execution, the robot must infer what has been completed, determine what remains, and maintain an appropriate subtask command. An incorrect subtask choice cannot generally be corrected by more precise motor control: the robot may execute the wrong subtask perfectly.

Vision-language-action (VLA) models provide a natural foundation for long-horizon robotic control by coupling the semantic representations of pretrained vision-language models with continuous action generation [5, 43, 19, 28, 3, 32]. Hierarchical VLAs further expose subtasks as an explicit interface for decomposing and executing multi-stage behaviors [18, 2, 35, 27, 13, 14, 32, 33]. Many hierarchical systems still make each high-level decision with a fixed inference budget, directly mapping the current observation and execution history to the next subtask. In this setting, the model neither compares alternative subtasks nor evaluates them through the physical states they are expected to produce. Recent systems instead use high-level tree search or recursive subgoal refinement [30, 40, 42], motivating the question we study here: how to couple open-ended language-subtask search with visual outcome prediction and execution memory in a generalist real-robot hierarchy. Without pre-commitment evaluation, an incorrect decision is typically detected only after execution has already altered the environment, at which point replanning can respond but cannot recover the cost of the failed commitment. Thus, the central bottleneck in long-horizon execution is no longer the availability of a hierarchical action interface, but the inference procedure used to generate the next subtask.

We instead formulate generating the next subtask as an inference-time reasoning problem, allowing computation to scale with the difficulty of each decision, as test-time computation does in language models [29, 9]. The subtask is a natural unit for this reasoning. Action-level search can resolve local control ambiguities, but often exposes only local physical consequences [26, 20, 8, 38, 15, 41]. Languageonly reasoning operates over longer horizons, but compares alternatives without grounding them in the physical states they would produce. Subtasks lie between these extremes: they are sparse enough to justify additional computation, yet temporally extended enough to induce meaningful and distinguishable changes in the environment. Comparing candidate subtasks before execution therefore requires predicting the state that each candidate would produce. Generative models of future observations provide precisely this capability [10, 4, 11, 12, 24], enabling candidate subtasks to be evaluated through their anticipated physical consequences before they are issued for execution.

We present τ₀-VLA, a hierarchical VLA system that instantiates this approach. At each inference step, a memoryaugmented high-level policy uses the current observation and execution memory to generate the appropriate subtask.

The proposal is used directly when the policy is confident. Otherwise, token-confidence statistics trigger an additional reasoning procedure. This procedure follows a propose– predict–evaluate loop. The proposal model generates openended candidate subtasks, a world model predicts the terminal observation induced by each candidate, and a value model scores candidate quality from the predicted outcomes. Beam search recursively expands promising candidates, allowing the inference budget to scale through the branching factor, beam width, and search depth. A reflective model then conditions on the retained branches and generates the final subtask. Its output may coincide with a proposal but is not restricted to the candidate set. The low-level policy then executes the generated subtask. The resulting real observation is incorporated into execution memory at the next inference step, closing the loop between predicted and observed consequences. To support reliable reasoning over extended executions, we train the memory mechanism to correct its own state. Specifically, we perturb memories derived from existing demonstrations and train the high-level policy to repair records that lag behind, run ahead of, or otherwise misrepresent the robot’s actual progress, requiring no additional annotation.

Once generated, each subtask is executed by a generalist low-level policy that combines a pretrained vision-language backbone with a Mixture-of-Transformers action expert. The policy operates over a unified 40-dimensional state and action space covering end effectors, arm joints, grippers, the waist, and the mobile base. This unified representation allows a single model to support fixed-base manipulation, bimanual coordination, and mobile whole-body control across multiple robot embodiments. We train the policy on approximately 40,000 hours of robot data collected from heterogeneous sources, together with multimodal co-training data, yielding a language-conditioned execution interface without the need for subtask-specific controllers.

We evaluate the complete system across multiple robot embodiments on real-world, long-horizon manipulation tasks. The evaluation covers room cleaning, meal preparation, tea making, and laundry collection, with episodes lasting up to 12 minutes. Hierarchical test-time computation substantially improves task success over whole-task inference using the same low-level policy and increases next-subtask prediction accuracy, with further gains obtained from larger inference budgets.

Our central contribution is a complete VLA system, τ₀- VLA, that formulates high-level subtask generation as a compute-scalable inference problem while retaining a unified low-level control interface across multiple robot embodiments. Our experiments show that allocating additional computation to high-level decisions substantially improves next-subtask prediction, and yields higher end-to-end task success on longhorizon real-robot tasks. We further provide a detailed empirical analysis of execution memory, consequence-aware search, and the compute-accuracy trade-off.

<!-- End of PDF page 2 -->

## II. RELATED WORK

τ₀-VLA brings together three lines of work: generalist VLA models for low-level control, hierarchical robot policies for subtask-level decision making, and world models with testtime computation for evaluating candidate decisions. We thus survey works in these areas.

## _A. Vision-Language-Action Models_

Vision-language-action (VLA) models map visual observations and language instructions to robot actions, often by adapting pretrained vision-language backbones for continuous control. Early work established scalable languageconditioned robot learning and transfer from vision-language pretraining [5, 43]. Subsequent work broadened this paradigm across tasks and embodiments [19, 28], while recent generalist models target continuous control, open-world generalization, and cross-embodiment transfer [3, 32, 27, 39, 1, 14]. Complementary work studies compact and efficient deployment [36]. Cross-embodiment policies must also reconcile heterogeneous control interfaces. RDT-1B introduces a physically interpretable unified action space, while Being-H0.5 maps heterogeneous controls into semantically aligned slots [22, 25].

This line of work primarily improves control representations, training scale, and transfer across tasks and embodiments. Our low-level policy follows the same generalist VLA paradigm. Our primary focus is complementary: making high-level subtask generation an explicit, compute-scalable inference problem.

## _B. Hierarchical Robot Policies_

Hierarchical policies separate temporally extended decisions from continuous control. Prior work grounds language-model decisions in learned robot affordances [18], represents action hierarchies through language [2], and connects slower semantic reasoning to faster visuomotor control through explicit commands, latent representations, or staged inference [35, 13, 27, 14, 32]. Recent systems address long-horizon state tracking and adaptation more explicitly. MemER retrieves task-relevant historical keyframes before generating instructions for a lowlevel VLA, while Goal2Skill combines structured memory, outcome verification, and reflection in a high-level VLM and low-level VLA hierarchy [37, 23]. Anticipation-VLA adaptively generates recursive subgoals, whereas VISTA uses a world model as the high-level policy to generate textual and visual subgoal sequences [42, 24].

Among generalist hierarchical foundation models, _π_ 0 _._ 7 also couples a semantic high-level policy, a world model, and a low-level policy [33]. Its world model generates subgoal images for a subtask already produced by the high-level policy. In τ₀-VLA, visual prediction instead occurs before commitment and supports comparison among multiple candidate subtasks through value-guided search and reflection. Thus, visual prediction conditions how a generated subtask is executed in _π_ 0 _._ 7, whereas it informs which subtask to execute in τ₀-VLA. Our high-level policy also maintains a correctable execution

memory and scales decision-time computation through search width and depth.

## _C. World Models and Test-Time Computation_

World models predict future states under candidate behavior and support planning and policy learning [16, 17]. In robot manipulation, generated images or videos have been used as goals for inverse dynamics and low-level control [10, 4], modeled jointly with actions [7, 6], and incorporated into search or iterative plan revision [11]. Reflective Planning iteratively revises a VLM plan using imagined future states [12]. Our search instead maintains multiple language-subtask branches and prunes them with a dedicated value model before reflective generation. At the action level, additional inference is allocated through sampling, verification, or value-guided selection [26, 20, 8]. FOREWARN predicts latent outcomes for low-level action plans and uses a VLM to select among them, while VLA-Reasoner performs world-model-guided Monte Carlo tree search over action trajectories [38, 15]. WLA-0 jointly predicts textual subtasks, future images, and actions, and its test-time scaling mode selects among action chunks using imagined future frames and a value model [41].

Search has also been applied to high-level robot decisions. VINE proposes subgoal transitions on a two-dimensional scene graph, predicts their success probabilities from successful and failed demonstrations, and prunes low-feasibility branches before execution [30]. Seeing Farther and Smarter combines predicted visual dynamics, value-guided beam search, confidence-based routing, and multi-path reflection over manipulation plans [40]. These are among the closest search-based mechanisms to our work. VINE searches structured scene-graph transitions using feasibility scores, while Seeing Farther and Smarter searches discrete manipulation plans with a learned critic. In contrast, τ₀-VLA searches open-ended language subtasks and evaluates each candidate through its predicted terminal image. It conditions subsequent decisions on a correctable execution memory, then uses a reflective model to generate the final subtask from the retained branches.

## III. PRELIMINARIES

We consider a robot task specified by a language instruction _ℓ_ . At inference step _t_ , a vision-language-action (VLA) policy maps the current multi-view observation _ot_ , proprioceptive state **s** _t_ , and language command _ct_ to an _H_ -step action chunk **a** _t_ : _t_ + _H−_ 1. It also receives textual control metadata _η_ specifying the embodiment, control mode, and whole-body configuration. Its concrete text serialization is provided in Appendix C-B:


![](assets/paper_tau0-vla-2026.pdf-0003-14.png)


where _θ_ denotes the policy parameters and _H_ is the actionchunk horizon. In direct execution, the full task instruction is used throughout the episode, so _ct_ = _ℓ_ . The policy must therefore infer the current stage of the task while generating the corresponding motor actions.

<!-- End of PDF page 3 -->

A hierarchical VLA separates high-level subtask generation from low-level action generation. At each inference step, the high-level policy first generates a subtask, and the lowlevel policy then generates actions conditioned on it. Let _µ_ denote the high-level policy, let _M_ 0 denote the initial empty execution memory, and let _z_ 0<sup>_⋆_=∅.Atinferencestep</sup><sup>_t_,</sup><sup>_µ_</sup> forms the high-level decision context


![](assets/paper_tau0-vla-2026.pdf-0004-01.png)


where _Mt−_ 1 is the carried execution memory and _zt_<sup>_⋆_</sup> _−_ 1<sup>isthe</sup> subtask generated at the preceding inference step. The policy brings the execution memory up to date and generates the subtask for the current observation:


![](assets/paper_tau0-vla-2026.pdf-0004-03.png)


Here _Mt_ summarizes the execution history observed up to _ot_ and therefore does not yet contain the outcome of executing _zt_<sup>_⋆_,while</sup><sup>_z_</sup> _t_<sup>_⋆_isthesubtaskgeneratedforthecurrentobserva-</sup> tion. The generated subtask is used as the language command, _ct_ = _zt_<sup>_⋆_,yielding</sup>


![](assets/paper_tau0-vla-2026.pdf-0004-05.png)


After executing the action chunk, the resulting real observation is used at the next inference step. At every inference step, the low-level policy conditions on the generated subtask _zt_<sup>_⋆_.Ifit</sup> remains the same as _zt_<sup>_⋆_</sup> _−_ 1<sup>,executionofthecurrentsubtask</sup> continues. If it changes, the low-level policy begins executing the newly generated subtask.

The high-level policy can either commit to its direct proposal or invoke additional test-time computation before producing _zt_<sup>_⋆_.Thus,thehigh-levelpolicydetermines</sup><sup>_which_</sup> _subtask is currently appropriate_ , while the low-level policy determines _how to execute it_ .

## IV. METHOD

## _A. System Overview_

τ₀-VLA consists of two components applied sequentially at each logical inference step. The _high-level policy_ maintains execution memory, decides when to invoke additional testtime computation, and generates the current subtask. The _lowlevel policy_ then maps the current observation and generated subtask to robot actions. This hierarchy supports long-horizon progress tracking and consequence-aware planning without modifying the low-level control interface. Figure 2 provides an overview. The following subsections define both policies and then describe their joint inference procedure.

## _B. High-Level Policy_

Unlike a conventional fixed-compute high-level policy which commits to a subtask through a single forward pass, our high-level policy _µ_ generates the next subtask with a budgetadaptive test-time computation (TTC) procedure. On uncertain predictions, TTC performs world-model-guided beam search over possible subtask sequences. It expands candidate branches, predicts and scores their visual outcomes, and passes

the retained beam to a reflective model that generates the final subtask. This procedure allows the policy to improve nextsubtask prediction by allocating additional test-time computation.

Accordingly, the high-level policy _µ_ consists of a proposal model _P_ , a world model _W_ , a value model _V_ , and a reflective model _F_ . Their inference interfaces are defined below.

**Proposal model.** At the beginning of inference step _t_ , the proposal model receives the context _ht_ defined in Section III, updates the execution memory, and generates a direct subtask proposal _zt_<sup>dir:</sup>


![](assets/paper_tau0-vla-2026.pdf-0004-16.png)


The proposal model updates the memory from _Mt−_ 1, the previously generated subtask _zt_<sup>_⋆_</sup> _−_ 1<sup>, and the current observation</sup> _ot_ . From token confidences produced by the same forward pass, the adaptive router computes _gt ∈{_ 0 _,_ 1 _}_ . Here _gt_ = 0 selects the fast route and _gt_ = 1 invokes TTC. The routing rule is defined in Appendix C-J.

**World model and value model.** For a generic candidate within beam search, let _o_ ˜ denote a head-camera RGB image and _z_ the candidate subtask. The world model predicts the terminal head-camera image, and the value model scores that outcome:


![](assets/paper_tau0-vla-2026.pdf-0004-19.png)


Thus, the world model always operates on a single headcamera image. The value model receives the global task instruction, candidate subtask, and predicted terminal image, and returns a scalar candidate-quality score _v_ .

**Test-time search.** When _gt_ = 1, the high-level policy invokes the SEARCH operation in Algorithm 1:


![](assets/paper_tau0-vla-2026.pdf-0004-22.png)


The operation performs beam search over candidate subtask sequences. The positive integers _N_ , _B_ , and _D_ denote the branching factor, beam width, and search depth. A branch _b_ stores a proposal context _h_ ( _b_ ), an ordered sequence of imagined subtasks _ρ_ ( _b_ ), a cumulative score _S_ ( _b_ ), and, for a non-root branch, a terminal predicted image _o_ ˆ( _b_ ). We initialize _Qt,_ 0 with a single root branch _b_ root whose context is _ht_ , whose path is the empty sequence _ρ_ ( _b_ root) = (), and whose score is _S_ ( _b_ root) = 0.

At each depth _d ∈{_ 1 _, . . . , D}_ , we independently invoke the proposal model _N_ times for every retained branch _b ∈Qt,d−_ 1:


![](assets/paper_tau0-vla-2026.pdf-0004-25.png)


Here _P_ ( _· | h_ ( _b_ )) denotes the proposal model’s decoding distribution. The sample index _i_ keeps repeated text proposals distinct and associates each proposal with its branch-local memory. For the root branch, _h_ ( _b_ root) = _ht_ , so the proposal model receives the current multi-view observation _ot_ . For every non-root branch, the visual component of _h_ ( _b_ ) is _o_ ˆ( _b_ ), the terminal head-camera image imagined for that branch. Routing decisions from proposal calls inside search are ignored.

<!-- End of PDF page 4 -->


![](assets/paper_tau0-vla-2026.pdf-0005-00.png)


<!-- Start of picture text -->
c. World-Model-Guided Test-Time Computation<br>Current context<br>observation<br>task Beam Search<br>memory Candidate Subtask Imagined Outcome Score<br>“Pick Up Spoon” 0.95<br>propose<br>High- World Value<br>Level  “Place Cup” Model Model 0.50 ×<br>Policy<br>“Pick Up Cup” 0.95<br>after depth D<br>reflect & commit<br>expand retained beams · repeat to depth D<br>Next Subtask Example Retained Branch:  Pick Up Spoon →Scoop Jelly →Pour Jelly<br>Pick Up Spoon<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0005-01.png)


<!-- Start of picture text -->
Observation Task Instruction Execution Memory<br>multi-view “Make Milk Tea” “Placed Cup”<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0005-02.png)


<!-- Start of picture text -->
Vision-Language Model<br>Qwen3.5-9B<br>Thought Updated Memory Candidate Subtask<br><think> “Placed Cup” “Pick Up Spoon”<br><!-- End of picture text -->


![](assets/paper_tau0-vla-2026.pdf-0005-03.png)


<!-- Start of picture text -->
Meta  Robot  Noised  Committed Subtask<br>Info State Action Pick Up Spoon<br>Vision-Language Model MoT<br>Qwen3.5-2B Action Expert<br><!-- End of picture text -->

Fig. 2: **The hierarchical** τ₀ **-VLA architecture.** (a) At a high-level inference step _t_ , the proposal model _P_ conditions on the latest multi-view observation _ot_ , task instruction _ℓ_ , carried execution memory _Mt−_ 1, and previously generated subtask _zt_<sup>_⋆_</sup> _−_ 1<sup>.</sup> It produces observation-aligned memory _Mt_ and a direct proposal _zt_<sup>dir.(b)Thelow-levelpolicyconditionsonthegenerated</sup> subtask _zt_<sup>_⋆_,multi-viewobservation</sup><sup>_ot_,proprioceptivestate</sup><sup>**s**</sup><sup>_t_,andtextualcontrolmetadata</sup><sup>_η_.Avision-languagebackboneand</sup> Mixture-of-Transformers (MoT) action expert generate the action chunk **a** _t_ : _t_ + _H−_ 1 through conditional flow matching from a noisy action chunk. (c) On the TTC route, the proposal model is invoked _N_ times for each retained branch to generate _N_ candidates. Given the branch’s head-camera image and a candidate, the world model predicts the terminal head-camera image, and the value model assigns a candidate-quality score conditioned on the task instruction, candidate, and predicted image. Beam search globally retains the top- _B_ branches by cumulative score and recursively expands them to depth _D_ . The figure illustrates the root expansion with _N_ = 3 and _B_ = 2, where local and cumulative scores coincide. Deeper expansion is omitted for clarity. The reflective model then conditions on _h_<sup>¯</sup> _t_ and the final branch summaries _Ct_ to generate the final subtask _zt_<sup>_⋆_.This</sup> output may coincide with a retained proposal but is not restricted to the retained set and is passed to the low-level policy for execution.

We write _b ⊕ z_ for the child branch obtained by appending subtask _z_ to branch _b_ . Its path _ρ_ ( _b⊕z_ ) is the ordered sequence _ρ_ ( _b_ ) followed by _z_ . Let _o_ ˜( _b_ ) be the head-camera image in _ot_ ˆ when _b_ = _b_ root and _o_ ( _b_ ) otherwise. For each indexed proposal, the world and value models compute


![](assets/paper_tau0-vla-2026.pdf-0005-06.png)


The child retains the corresponding memory _M_<sup>_b,i_</sup> _t,d_<sup>andpre-</sup> dicted image. Its cumulative score and context for the next expansion are


![](assets/paper_tau0-vla-2026.pdf-0005-08.png)


All memories produced inside search are branch-local and never overwrite the persistent execution memory _Mt_ .

At depth _d_ , the indexed collection of all children is


![](assets/paper_tau0-vla-2026.pdf-0005-11.png)


Before pruning, _At,_ 1 contains _N_ children, while each subsequent expansion produces at most _BN_ children. The top- _B_ operation ranks all children globally by cumulative score. Each retained child’s predicted image and branch-local memory define the context for its next expansion. This process repeats until depth _D_ .

The final beam is summarized by the indexed collection


![](assets/paper_tau0-vla-2026.pdf-0005-14.png)


where _ρ_ ( _b_ ) is the imagined subtask path and ˆ _o_ ( _b_ ) is its terminal predicted head-camera image.

**Reflective model.** At the end of the TTC route at inference step _t_ , the reflective model conditions on the retained branch summaries _Ct_ and the observation-aligned real context _h_<sup>¯</sup> _t_ = ( _ℓ, Mt, zt_<sup>_⋆_</sup> _−_ 1<sup>_, ot_).Itgeneratesthefinalsubtaskpassedtothe</sup> low-level policy:


![](assets/paper_tau0-vla-2026.pdf-0005-17.png)


Its output may reproduce a retained proposal but is not constrained to the candidate set. The reflective model does not update persistent execution memory.

<!-- End of PDF page 5 -->

## _C. Low-Level Policy_

The low-level policy couples a vision-language backbone with a Mixture-of-Transformers (MoT) action expert. At each full-attention layer, the action tokens and backbone tokens interact through joint attention while being processed by separately parameterized Transformer streams. The policy conditions on multi-view observations, proprioceptive state, and a language command. The action expert is trained with conditional flow matching to learn a velocity field from noise to the distribution of action chunks. At inference, this field is integrated to generate executable actions.

**Unified state and action space.** We represent heterogeneous embodiments in a shared 40-dimensional state and action space. Per-sample state and action masks select the valid channels, while the control metadata specifies the control parameterization. This interface allows the same policy to support fixed-base and whole-body control without embodimentspecific output heads. Appendix C-B provides the complete layout and action encoding.

**Masked flow matching.** Because different embodiments occupy different channels in the unified action space, we mask both the flow path and its training objective. Let _da_ denote the action dimension and let **M** _∈{_ 0 _,_ 1 _}_<sup>_da×da_</sup> be a diagonal action-mask matrix, where [ **M** ] _ii_ = 1 if channel _i_ is active and [ **M** ] _ii_ = 0 otherwise. The same matrix is applied to every action vector in the chunk. For each _j ∈{_ 0 _, . . . , H −_ 1 _}_ , let **a** _t_ + _j,_ **_ϵ_** _t_ + _j ∈_ R<sup>_da_</sup> , where **_ϵ_** _t_ + _j ∼N_ ( **0** _,_ **I** _da_ ) and **I** _da_ is the _da_ -dimensional identity matrix. We use the reverse-time parameterization of the linear flow path in _π_ 0 [3], matching the task-specific checkpoints evaluated in this work. A flow time _τ ∈_ [0 _,_ 1] is shared across the chunk, with _τ_ = 1 denoting noise and _τ_ = 0 denoting a clean action. The interpolated action and target velocity are


![](assets/paper_tau0-vla-2026.pdf-0006-04.png)


Given the complete noisy chunk, the action expert jointly predicts the velocities of all _H_ action tokens. We supervise only active channels and project inactive channels before each velocity-field evaluation and at the final output. Appendix C-B provides the full loss and sampling details.

## _D. System Inference_

At inference, the high-level policy and low-level policy form a closed loop. At each logical inference step, the proposal model updates the execution memory and generates _zt_<sup>dir.</sup> When _gt_ = 0, this direct proposal becomes the final subtask _zt_<sup>_⋆_.When</sup><sup>_gt_=1,test-timesearchproduces</sup><sup>_Ct_,andthe</sup> reflective model generates _zt_<sup>_⋆_fromtheobservation-aligned</sup> context and retained branch summaries. The low-level policy then generates and executes an action chunk conditioned on _zt_<sup>_⋆_. The resulting real observation is incorporated into memory</sup> at the next inference step.

Algorithm 1 presents these logical dependencies as a sequential procedure. In deployment, the high-level and lowlevel policies are pipelined asynchronously, as detailed in

|**Alg**|**orithm 1** Closed-loop system inference|
|---|---|
|**Re**|**quire:** Task instruction _ℓ_, models _P, W, V, F, πθ_|
|**Re**|**quire:** Control metadata _η_, action horizon _H_, branching<br>factor _N_, beam width _B_, depth _D_|
|1:|Initialize _M_0 _←_∅and _z_<sup>_⋆_</sup><br>0 <sup>_←_∅</sup>|
|2:|**for** _t_= 1_,_2_, . . ._ until task termination **do**|
|3:|Observe _ot_ and **s**_t_<br><br>|
|4:|_ht ←_(_ℓ, Mt−_1_, z_<sup>_⋆_</sup><br>_t−_1<sup>_, ot_)</sup><br><br>|
|5:|(_z_<sup>dir</sup><br>_t _<sup>_, Mt_)</sup><sup>_←P_(</sup><sup>_ht_)</sup>|
|6:|Compute _gt_ from the proposal token confdences<br>¯<br>|
|7:|_ht ←_(_ℓ, Mt, z_<sup>_⋆_</sup><br>_t−_1<sup>_, ot_)</sup><br>|
|8:|**if** _gt_ = 0 **then**<br>|
|9:|_z_<sup>_⋆_</sup><br>_t _<sup>_←z_dir</sup><br>_t_<br>_▷_fast route|
|10:|**else**<br>|
|11:|_Ct ←_SEARCH(_ht, P, W, V, N, B, D_)<br><sup>¯</sup>|
|12:|_z_<sup>_⋆_</sup><br>_t _<sup>_←F_(</sup><sup>_ht, Ct_)</sup>|
|13:|**end if**|
|14:|**a**_t_:_t_+_H−_1 _←πθ_(_ot,_**s**_t, z_<sup>_⋆_</sup><br>_t _<sup>_, η_)</sup>|
|15:<br>16:|Execute action chunk **a**_t_:_t_+_H−_1<br> **end for**|



Appendix C-F. The same routing statistics are used across tasks, with task-specific thresholds calibrated on held-out data. In direct-execution mode, the system bypasses the high-level policy and conditions the low-level policy directly on _ℓ_ .

## V. TRAINING

## _A. Data Sources_

**Low-level policy data.** The low-level policy is trained on 40,115 hours of heterogeneous real-world data spanning fixedbase, mobile, and bimanual embodiments. The corpus combines human demonstrations, autonomous policy rollouts, and UMI-style recordings, covering manipulation, navigation, and whole-body coordination across diverse robot morphologies and control interfaces. We interleave these trajectories with multimodal vision-language data for instruction following, visual grounding, spatial and depth reasoning, and robot-centric perception. This co-training mixture preserves the semantic and visual capabilities of the VLM backbone during action learning. Appendix C-D provides further details on the data composition and processing.

**High-level policy data.** Supervision for the high-level policy is derived automatically from existing task, stage, and executable-subtask annotations together with segmented multiview demonstrations. These sources supervise reasoning over task progress, execution-memory updates, and subtask generation. We additionally construct memory-perturbed examples that teach the policy to recover when its execution history lags behind, runs ahead of, or otherwise conflicts with the observed state. Subtask-aligned start–end frames supervise the world model. This construction requires no additional per-sample human labeling. Appendix C-E describes the annotation, filtering, and quality control pipeline.

<!-- End of PDF page 6 -->

## _B. Low-Level Policy Training_

We initialize the low-level vision-language backbone from Qwen3.5-2B [34] and attach the MoT action expert described in Section IV-C. The resulting policy maps the current observation and language command to robot actions in the unified action space, as defined in Eq. (1). We train it in three successive stages.

_a) Stage 1: Knowledge-isolated co-training.:_ We jointly train on multimodal and robot-action data with knowledge isolation (KI) [31]. The MoT action expert attends to the visionlanguage backbone, while action-loss gradients are blocked at the backbone interface. This allows the action expert to learn control-relevant representations without prematurely perturbing the pretrained backbone, with multimodal supervision stabilizing training.

_b) Stage 2: End-to-end co-training.:_ We remove KI and optimize the full model end to end, retaining multimodal data as auxiliary supervision. This allows action gradients to adapt the backbone toward representations better suited for robot control and strengthens the coupling between perception, language, and action prediction.

_c) Stage 3: Task-specific adaptation.:_ For each deployment task, we fine-tune the policy on a small set of taskspecific demonstrations, adapting it to the target embodiment, camera viewpoints, object configurations, and success criteria.

to _clearly correct_ , mapped to scalar values in [0 _._ 05 _,_ 0 _._ 95]. Training pairs each sampled state with a candidate subtask from offline rollouts generated by repeatedly applying the proposal and world models. The candidates are graded against the ground-truth next step, aligning training with the proposals that the value model scores during search.

**Reflective model.** We train the reflective model with the same observation-aligned context and retained-branch summaries used at inference. The candidate plans come from the same offline rollouts together with their predicted future states and scores. Each set is paired with the ground-truth next subtask as an autoregressive target. This supervision teaches the model to ground its generation in visual evidence and repair flawed candidate plans rather than simply copy the first proposal.

## VI. EXPERIMENTS

Our evaluation separates high-level decision quality, longhorizon system performance, and short-horizon execution across embodiments. At the high level, we evaluate nextsubtask prediction, test-time computation, and the contribution of execution memory. We then evaluate the hierarchical system on four long-horizon tasks, study test-time computation on Book Organization task, and evaluate the VLA through direct execution on the shorter Collect Laundry and Tidy Makeup Table tasks.

## _C. High-Level Policy Training_

The proposal, value, and reflective models are independently fine-tuned from the same robot-pretrained VLM checkpoint initialized from Qwen3.5-9B [34]. The world model is initialized from Step1X-Edit [21] and trained separately. We describe the specific supervision for each model below.

**Proposal model.** The proposal model generates candidate subtasks while maintaining and updating execution memory. We train it on both aligned execution histories and automatically perturbed histories that lag behind, run ahead of, or misrepresent progress after a failure. The corrected targets teach the model to reconcile memory with visual evidence before generating the next subtask. The perturbation types, recovery targets, and sampling mixture are detailed in Appendix C-E. **World model.** The world model is trained to predict the visual outcome of a candidate subtask at completion. Each training example pairs the head-camera RGB observations at the beginning and end of an annotated subtask segment, conditioned on the corresponding subtask instruction. These subtask-aligned transitions are drawn from real trajectories.

Starting from real observations, we construct multi-step simulated rollouts by recursively alternating subtask proposal with _P_ and visual outcome prediction with _W_ , using each predicted observation to condition the next proposal step. These simulated rollouts form an offline dataset used to train the value and reflective models described below.

**Value model.** We formulate value prediction as a multiplechoice VQA task. Conditioned on the global instruction, candidate subtask, and its imagined terminal image, the model predicts one of five ordinal quality levels, from _clearly wrong_

## _A. Evaluation Setup_

**Robot platforms.** We evaluate on AGIBOT G1 for the four primary long-horizon tasks, ARX AC One for Book Organization and Collect Laundry, and a bimanual Franka Research 3 setup for Tidy Makeup Table. All platforms provide multiview RGB observations and proprioceptive state and map their controls into the unified action space. Hardware and sensing details are deferred to Appendix C-A.

**Task suite.** The evaluation covers six household tasks and one benchmark comprising three separately evaluated instruction-following task groups. Clean Room, Prepare Ingredients, Tomato and Egg Stir Fry, and Collect Laundry require mobile manipulation. Make Milk Tea, Book Organization, and the three Tidy Makeup Table groups require manipulation without base motion. We use Tidy Makeup Table for instruction-following evaluation. Prepare Ingredients and Tomato and Egg Stir Fry form a collaborative cooking workflow for the same tomato and egg dish. One robot prepares the ingredients, and another completes the cooking stage. We therefore evaluate and report the two stages as separate tasks.

- **Clean Room (25 steps).** The robot enters the bedroom, retrieves two dirty garments from the nightstand and bed, and places them in a laundry basket. It then hangs a handbag on a clothes rack, retrieves a blanket and hands it to a person, leaves the bedroom, and disposes of snackbag trash from the coffee table. A typical successful rollout lasts approximately 8 min.

- **Prepare Ingredients (14 steps).** The robot moves between the refrigerator and preparation table, opens the

<!-- End of PDF page 7 -->


![](assets/paper_tau0-vla-2026.pdf-0008-00.png)


<!-- Start of picture text -->
(a) Clean Room<br>(b) Prepare Ingredients<br>(c) Tomato and Egg Stir Fry<br>(d) Make Milk Tea<br>(e) Collect Laundry (f) Tidy Makeup Table<br><!-- End of picture text -->

Fig. 3: **Representative physical-robot evaluation tasks.** (a) Clean Room requires collecting two dirty garments, placing them in a laundry basket, hanging a handbag, handing a blanket to a person, and disposing of table trash. (b) Prepare Ingredients requires retrieving a tomato and an egg, then cracking and stirring the egg while returning the tools. (c) Tomato and Egg Stir Fry requires chopping and transferring the tomatoes, cooking and seasoning the ingredients, plating the dish, and returning the cookware and utensils. (d) Make Milk Tea requires adding toppings, pouring milk and tea, sealing the cup, and inserting a straw. (e) Collect Laundry requires transferring a T-shirt from the bedside table to a laundry basket. (f) Tidy Makeup Table comprises three separately scored instruction-following groups that require different object selections and action sequences from matched visual states.

   - refrigerator, retrieves a tomato and an egg, and closes the refrigerator. It then places the ingredients in bowls, cracks the egg with an egg cracker, stirs the egg mixture, and returns both tools to a porcelain plate. A typical successful rollout lasts approximately 4 min.

- **Tomato and Egg Stir Fry (22 steps).** The robot chops the tomatoes and transfers the tomatoes and prepared egg mixture to the induction cooker. It turns on the cooker, adds oil and the egg mixture, stir-fries the eggs, adds the tomatoes and salt, finishes stir-frying, plates the dish, returns the cookware and utensils, turns off the cooker, places the finished dish on the preparation table, and returns both arms to the home position. A typical successful rollout lasts approximately 10 min.

- **Make Milk Tea (13 steps).** The robot places a cup at the preparation station, adds two toppings, pours milk and tea in sequence, seals the cup with a lid, and inserts a straw.

A typical successful rollout lasts approximately 3 min.

- **Collect Laundry (5 steps).** The robot searches for and moves to the drawer cabinet, picks up the dirty T- shirt from the bedside table with its left arm, closes the cabinet door with its right arm, searches for and moves to the laundry basket, and places the T-shirt in the basket with its left arm. A typical successful rollout lasts approximately 1 min.

- **Tidy Makeup Table (three task groups).** This benchmark pairs matched visual states with instructions that specify different target objects, active arms, action sequences, and destinations. We score three groups independently. In _Cotton Pad_ (2 steps), the left arm picks up the cotton pad and places it in its designated compartment. In _Eyelash Curler_ (2 steps), the left arm picks up the eyelash curler and places it in its designated compartment. In _Makeup Puff_ (4 steps), both arms open

<!-- End of PDF page 8 -->

the drawer, the left arm picks up the makeup puff from the tabletop and places it inside, and both arms close the drawer. Executing all eight steps across the three task groups takes approximately 30 s in total.

- **Book Organization (3 steps).** The robot uses its two arms to reorder four books of different heights across four fixed slots, swapping any pair of books at each step until they are arranged from tallest to shortest. A typical successful rollout lasts approximately 1 min. We evaluate two initial-state regimes while keeping the task objective and action space fixed: in-domain initial arrangements appear in the training data, whereas out-of-domain (OOD) initial arrangements are not observed during training. The OOD setting therefore evaluates the high-level policy’s ability to predict appropriate swaps from observations induced by unseen initial configurations.

**Evaluation protocols.** We use complementary protocols for high-level decision making and physical-robot execution. _High-level policy evaluation._ We isolate high-level decision making on annotated subtask-boundary samples and report next-subtask prediction accuracy and inference time. All methods within each comparison receive the same inputs and output format. Evaluation-set construction, variant-specific inputs, and judging details are provided in Appendix C-H.

_Physical-robot evaluation._ We report _success rate_ (SR) and _progress_ using the same task-specific terminal conditions for all methods. SR measures the fraction of successful trials. A trial is successful only when every required milestone is completed. Skipping a required milestone makes the trial unsuccessful. Progress measures how much of the task is completed using annotated milestones. Each milestone receives full, half, or zero credit based on completion quality, and dependent milestones count only after their required earlier steps are completed. Exact scoring rules, trial durations, reset procedures, and termination conditions are provided in Appendix C-C.

## _B. Overall System Evaluation_

Table I reports the four long-horizon tasks: Clean Room, Prepare Ingredients, Tomato and Egg Stir Fry, and Make Milk Tea. These tasks require navigation, object search, manipulation, state tracking, and recovery across multiple stages. For the controlled comparison of our two variants, the observation and action interfaces and the low-level policy are fixed. Standalone τ₀-VLA conditions the low-level policy on the full task instruction. The hierarchical system instead supplies bounded subtasks from the high-level policy. Both variants in this comparison run without beam search, which is evaluated separately in the test-time computation experiments. **Clean Room.** Performance on Clean Room benefits clearly from explicit execution memory. The hierarchical system retains progress across room transitions, whereas failures without this memory concentrate on handbag hanging and the later room-tidying stages.

**Prepare Ingredients.** Most failures in Prepare Ingredients trials arise during egg pickup, cracking, or stirring. Because

these actions are prerequisites for later preparation stages, an early error blocks substantial downstream progress. Tracking completed stages is therefore particularly valuable on this task. **Tomato and Egg Stir Fry.** The decisive bottleneck in Tomato and Egg Stir Fry is seasoning. Adding salt causes little visible change, so the current observation alone cannot reliably reveal whether this step has already been completed. Direct-execution policies therefore tend to add salt repeatedly or skip it entirely, either of which violates the success criterion. The hierarchical system resolves this ambiguity by recording seasoning progress explicitly.

**Make Milk Tea.** Make Milk Tea presents a complementary regime in which both τ₀-VLA variants already complete most of the sequence reliably. Both achieve an SR of 5/10 and more than 91% progress, outperforming the external baselines. Their remaining failures occur during lid attachment or straw insertion, after the preceding preparation stages have been completed. This pattern identifies final contact-rich manipulation as the main remaining bottleneck. When TTC is enabled, the hierarchical policy further improves to an SR of 7/10 and 95.38% progress, as reported in Table III.

## _C. Evaluation Across Embodiments_

We evaluate the low-level policy on two additional embodiments using Collect Laundry and Tidy Makeup Table. Because these tasks contain only two to five steps, each method executes the full task instruction without task decomposition, execution memory, or test-time search. This setting evaluates low-level control and language-conditioned manipulation across embodiments, separately from the high-level policy used in the long-horizon tasks. Table II reports the results.

**Collect Laundry.** Collect Laundry evaluates mobile manipulation by the low-level policy. The observed failures concentrate on T-shirt grasping and navigation around the bed: LingBot-VLA sometimes fails to lift the T-shirt, while GR00T collides with the foot of the bed and terminates the rollout. **Tidy Makeup Table.** The Tidy Makeup Table benchmark evaluates instruction following under matched visual states. The language instruction changes the target object, active arm, action order, and destination, so the observation alone does not determine the correct action sequence. Cotton Pad and Eyelash Curler emphasize instruction-conditioned object selection and placement, whereas Makeup Puff adds bimanual drawer manipulation. GR00T occasionally pauses during drawer motion or selects the makeup puff instead of the eyelash curler. LingBot-VLA generally follows the specified arm but tends to select visually nearby objects rather than the instructed target. _π_ 0 _._ 5 sometimes releases and re-grasps the makeup puff, closes the drawer less smoothly, or omits the cotton pad. These failures distinguish language-grounding errors from low-level execution inefficiencies: selecting or omitting the wrong object changes task completion, whereas re-grasping and hesitant drawer motion reduce execution efficiency even when the intended final state is reached.

<!-- End of PDF page 9 -->

TABLE I: **Long-horizon task performance.** Each method–task setting uses 10 independently collected physical-robot trials. SR reports successful trials as _x/_ 10, and Progress is the normalized milestone-completion score. Avg. is the unweighted mean of the four task-level rates. The first four rows use direct execution. The final row uses the Hierarchical System with Plan Once and no beam search.

|Method|Cle|an Room|Prepar|e Ingredients|Tomato|and Egg Stir Fry|Make|Milk Tea||Avg.|
|---|---|---|---|---|---|---|---|---|---|---|
||SR _↑_|Progress _↑_|SR _↑_|Progress _↑_|SR _↑_|Progress _↑_|SR _↑_|Progress _↑_|SR _↑_|Progress _↑_|
|GR00T N1.7 [27]|0/10|59.80%|1/10|68.57%|0/10|24.32%|0/10|28.46%|2.50%|45.29%|
|LingBot-VLA [1]|0/10|66.60%|0/10|35.00%|0/10|12.27%|0/10|63.85%|0.00%|44.43%|
|_π_0_._5 [32]|4/10|86.20%|2/10|73.93%|0/10|49.77%|3/10|82.31%|22.50%|73.05%|
|τ₀-VLA|4/10|92.80%|2/10|66.43%|0/10|65.00%|**5/10**|**96.15%**|27.50%|80.10%|
|τ₀-VLA (Hierarchical System, Plan Once)|**5/10**|**94.80%**|**4/10**|**82.86%**|**4/10**|**81.82%**|**5/10**|91.92%|**45.00%**|**87.85%**|



TABLE II: **Direct-execution performance across embodiments.** Tidy Makeup Table comprises three independently scored instruction-following groups. All methods execute the full task instruction without a high-level policy. Each task or task group is evaluated over 10 trials. SR reports successful trials as _x/_ 10, and Progress is the normalized milestone-completion score.

|Method|Collect|Laundry|||Tidy Ma|keup Table|||
|---|---|---|---|---|---|---|---|---|
||T-|shirt|Cott|on Pad|Eyelas|h Curler|Make|up Puff|
||SR _↑_|Progress _↑_|SR _↑_|Progress _↑_|SR _↑_|Progress _↑_|SR _↑_|Progress _↑_|
|GR00T N1.7 [27]|4/10|76.00%|**10/10**|87.50%|8/10|77.50%|7/10|52.50%|
|LingBot-VLA [1]|2/10|35.00%|9/10|67.50%|3/10|22.50%|3/10|33.75%|
|_π_0_._5 [32]|9/10|88.00%|9/10|85.00%|8/10|85.00%|7/10|73.75%|
|τ₀-VLA|**10/10**|**97.00%**|**10/10**|**95.00%**|**9/10**|**92.50%**|**10/10**|**95.00%**|



## _D. Test-Time Computation Experiments_

We conduct several experiments to quantify the gains in subtask-prediction accuracy achieved by test-time computation (TTC) and assess whether these gains lead to higher success rates in closed-loop real-robot evaluation. Our TTC experiments cover the following tasks: Make Milk Tea, Book Organization, and Clean Room. We evaluate Book Organization both in-domain and out-of-domain (OOD) initial book arrangements: the former appears in the training data, whereas the latter does not. The OOD setting evaluates the high-level policy’s performance when planning from unseen initial observations. We first measure next-subtask prediction accuracy in an open-loop setting, then evaluate TTC in closed-loop realrobot execution, and finally examine the relationship between computational cost and prediction accuracy.

TABLE III: **Closed-loop physical-robot performance with test-time computation.** Each entry uses 10 independently collected trials. Book Organization uses shuffled initial arrangements and is reported without the in-domain and OOD split used in the open-loop evaluation. SR denotes task success rate, and Progress is the normalized milestone-completion score.

|Method|Make|Milk Tea|Book|Organization|Cle|an Room|
|---|---|---|---|---|---|---|
||SR _↑_|Progress _↑_|SR _↑_|Progress _↑_|SR _↑_|Progress _↑_|
|Plan Once|5/10|91.92%|6/10|66.67%|5/10|94.80%|
|TTC|**7/10**|**95.38%**|**9/10**|**93.33%**|**7/10**|**97.60%**|




![](assets/paper_tau0-vla-2026.pdf-0010-08.png)


<!-- Start of picture text -->
Plan Once Best-of- N TTC (Ours)<br>90 87.3 88.0 87.0<br>83.0<br>80<br>74.0 74.0<br>72.0<br>70.0<br>70 64.7 66.0<br>60 57.5<br>50.0<br>50<br>40<br>Make Milk Tea Book Organization Book Organization Clean Room<br>(In-Domain) (OOD)<br>Accuracy (%)<br><!-- End of picture text -->

Fig. 4: **Next-subtask prediction accuracy under different high-level inference methods.** We compare Plan Once, Bestof- _N_ , and TTC (Ours) across four evaluation settings: Make Milk Tea, Book Organization (In-Domain), Book Organization (OOD), and Clean Room.

**Open-loop Subtask Prediction.** Following the high-level policy evaluation protocol above, we evaluate next-subtask prediction at different stages of each task. We compare TTC against two baselines. **Plan Once** makes one high-level prediction at each decision point without test-time search. It remains a variant of the high-level policy and is distinct from direct execution, which bypasses the high-level policy. **Best-of-** _N_ samples _N_ candidates, predicts the next observation for each candidate using the same world model as TTC, and selects the candidate assigned the highest score by the same value model.

As shown in Figure 4, TTC achieves the highest accu-

<!-- End of PDF page 10 -->


![](assets/paper_tau0-vla-2026.pdf-0011-00.png)


<!-- Start of picture text -->
(a) Make Milk Tea (b) Book Organization<br>90<br>80<br>85<br>75<br>80<br>70<br>75<br>65<br>70<br>60<br>Plan Once (64.7%) Plan Once (55.3%)<br>65 55<br>0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5<br>Test-Time Compute (PFLOPs / sample) Test-Time Compute (PFLOPs / sample)<br>TTC Result Fitted Curve Plan Once<br>Accuracy (%) Accuracy (%)<br><!-- End of picture text -->

Fig. 5: **Relationship between computational cost and subtask-prediction accuracy.** The accuracy increases with additional computation for the Make Milk Tea and Book Organization tasks, shown in panels (a) and (b), respectively. Each point represents an experimental result obtained at a different computational cost. The orange dashed curves show saturation fits to these results, and the gray dashed lines indicate the Plan Once baselines.

racy across all four evaluation settings, demonstrating its effectiveness in improving next-subtask prediction. Take the OOD Book Organization setting as an example, where TTC achieves 74 _._ 0% accuracy, compared with 50 _._ 0% for Plan Once and 57 _._ 5% for Best-of- _N_ . In the OOD setting, the initial book arrangements are absent from the training data, placing the corresponding observations outside the high-level policy’s training distribution. Under this distribution shift, directly predicting the next subtask with the fine-tuned VLM is more prone to error. Instead, TTC recursively expands candidate branches by using the world model to predict future observations and the value model to score the resulting imagined branches. The reflective model then uses the retained branches as context to generate the current subtask. Unlike direct VLM prediction based solely on patterns learned during fine-tuning, TTC evaluates candidate consequences at decision time, resulting in improved next-subtask prediction accuracy. Although Best-of- _N_ evaluates sampled candidates using the same world and value models, it performs only one-step selection without TTC’s multi-step expansion and reflective commitment. Consequently, its accuracy gains relative to Plan Once are consistently smaller than those achieved by TTC across the evaluated settings.

**Closed-loop Real-robot Evaluation.** We further assess whether TTC leads to higher task success rates in closed-loop real-robot evaluation. We deploy the high-level policy with and without TTC while keeping the low-level policy fixed. As reported in Table III, TTC improves both task progress and success rate across all evaluated tasks. This benefit is particularly important for tasks such as Book Organization, which do not admit a fixed execution plan. These results show that the gains in next-subtask prediction accuracy achieved by TTC lead to more effective and efficient closed-loop task execution.

**Computational Cost vs. Subtask-Prediction Accuracy.**

Finally, we conduct an analysis of the relationship between test-time computational cost and subtask-prediction accuracy. Figure 5 reports this relationship for the Make Milk Tea and Book Organization tasks. Each point represents an experimental result with a different computational cost, and the orange dashed curves provide approximate saturation fits to the observations. Specifically, we fit a saturating exponential function. The parameters are estimated by least squares using all experimental observations. Accuracy rises rapidly when computation is increased from a small budget, showing that additional search and reflection substantially improve subtask prediction in the low-compute regime. The marginal gain then gradually decreases and the fitted curves approach a plateau. This trend indicates that TTC provides a favorable compute– accuracy trade-off at moderate budgets, while its benefit eventually saturates as additional computation is allocated.

## VII. CONCLUSION

We present τ₀-VLA, a hierarchical VLA system for longhorizon manipulation that combines a memory-augmented high-level policy with a low-level policy. The high-level policy handles routine decisions efficiently and allocates worldmodel-guided computation when a decision benefits from explicit consequence evaluation. The low-level policy is trained from heterogeneous robot and vision-language data for deployable whole-body mobile manipulation. Together, the two levels provide a practical way to allocate reasoning over long task horizons while retaining stable real-robot execution.

## REFERENCES

- [1] Ant Group. LingBot-VLA: A pragmatic vla foundation model, 2026.

- [2] Suneel Belkhale, Tianli Ding, Ted Xiao, Pierre Sermanet, Quon Vuong, Jonathan Tompson, Yevgen Chebotar, Debidatta Dwibedi, and Dorsa Sadigh. RT-H: Action hierar-

<!-- End of PDF page 11 -->

chies using language. _arXiv preprint arXiv:2403.01823_ , 2024.

- [3] Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al. _π_ 0: A visionlanguage-action flow model for general robot control. _arXiv preprint arXiv:2410.24164_ , 2024.

- [4] Kevin Black, Mitsuhiko Nakamoto, Pranav Atreya, Homer Walke, Chelsea Finn, Aviral Kumar, and Sergey Levine. Zero-shot robotic manipulation with pretrained image-editing diffusion models. In _International Conference on Learning Representations_ , volume 2024, pages 33431–33452, 2024.

- [5] Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. RT-1: Robotics transformer for real-world control at scale. _arXiv preprint arXiv:2212.06817_ , 2022.

- [6] Jun Cen, Chaohui Yu, Hangjie Yuan, Yuming Jiang, Siteng Huang, Jiayan Guo, Xin Li, Yibing Song, Hao Luo, Fan Wang, et al. WorldVLA: Towards autoregressive action world model. _arXiv preprint arXiv:2506.21539_ , 2025.

- [7] Chi-Lam Cheang, Guangzeng Chen, Ya Jing, Tao Kong, Hang Li, Yifeng Li, Yuxiao Liu, Hongtao Wu, Jiafeng Xu, Yichu Yang, et al. GR-2: A generative videolanguage-action model with web-scale knowledge for robot manipulation. _arXiv preprint arXiv:2410.06158_ , 2024.

- [8] Mingtong Dai, Lingbo Liu, Yongjie Bai, Yang Liu, Zhouxia Wang, Rui Su, Chunjie Chen, Liang Lin, and Xinyu Wu. RoVer: Robot reward model as test-time verifier for vision-language-action model. _arXiv preprint arXiv:2510.10975_ , 2025.

- [9] DeepSeek-AI. Deepseek-r1: Incentivizing reasoning capability in LLMs via reinforcement learning. _arXiv preprint arXiv:2501.12948_ , 2025.

- [10] Yilun Du, Sherry Yang, Bo Dai, Hanjun Dai, Ofir Nachum, Josh Tenenbaum, Dale Schuurmans, and Pieter Abbeel. Learning universal policies via text-guided video generation. _Advances in Neural Information Processing Systems_ , 36:9156–9172, 2023.

- [11] Yilun Du, Sherry Yang, Pete Florence, Fei Xia, Ayzaan Wahid, Pierre Sermanet, Tianhe Yu, Pieter Abbeel, Joshua B Tenenbaum, Leslie Kaelbling, et al. Video language planning. In _International Conference on Learning Representations_ , volume 2024, pages 31138– 31155, 2024.

- [12] Yunhai Feng, Jiaming Han, Zhuoran Yang, Xiangyu Yue, Sergey Levine, and Jianlan Luo. Reflective planning: Vision-language models for multi-stage long-horizon robotic manipulation. In Joseph Lim, Shuran Song, and Hae-Won Park, editors, _Proceedings of The 9th Conference on Robot Learning_ , volume 305 of _Proceedings of Machine Learning Research_ , pages 2038–2062. PMLR, 2025. URL https://proceedings.mlr.press/v305/feng25b.

html.

- [13] Figure AI. Helix: A vision-language-action model for generalist humanoid control. https://www.figure.ai/news/ helix, 2025.

- [14] Google DeepMind. Gemini Robotics 1.5: Pushing the frontier of generalist robots with advanced embodied reasoning, thinking, and motion transfer, 2025.

- [15] Wenkai Guo, Guanxing Lu, Haoyuan Deng, Zhenyu Wu, Yansong Tang, and Ziwei Wang. VLA-Reasoner: Empowering vision-language-action models with reasoning via online Monte Carlo tree search. In _2026 IEEE International Conference on Robotics and Automation (ICRA)_ , 2026.

- [16] David Ha and J¨urgen Schmidhuber. World models. _arXiv preprint arXiv:1803.10122_ , 2018.

- [17] Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy P. Lillicrap. Mastering diverse control tasks through world models. _Nature_ , 640(8059):647–653, 2025. doi: 10.1038/s41586-025-08744-2. URL https://doi.org/10. 1038/s41586-025-08744-2.

- [18] Brian Ichter, Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, et al. Do as i can, not as i say: Grounding language in robotic affordances. In _Conference on Robot Learning_ , pages 287–318. PMLR, 2023.

- [19] Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. OpenVLA: An open-source vision-language-action model. _arXiv preprint arXiv:2406.09246_ , 2024.

- [20] Jacky Kwok, Christopher Agia, Rohan Sinha, Matt Foutter, Shulu Li, Ion Stoica, Azalia Mirhoseini, and Marco Pavone. RoboMonkey: Scaling test-time sampling and verification for vision-language-action models. In Joseph Lim, Shuran Song, and Hae-Won Park, editors, _Proceedings of The 9th Conference on Robot Learning_ , volume 305 of _Proceedings of Machine Learning Research_ , pages 3200–3217. PMLR, 2025. URL https://proceedings.mlr. press/v305/kwok25a.html.

- [21] Shiyu Liu, Yucheng Han, Peng Xing, Fukun Yin, Rui Wang, Wei Cheng, Jiaqi Liao, Yingming Wang, Honghao Fu, Chunrui Han, et al. Step1X-Edit: A practical framework for general image editing. _arXiv preprint arXiv:2504.17761_ , 2025.

- [22] Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, and Jun Zhu. RDT-1B: A diffusion foundation model for bimanual manipulation. In _The Thirteenth International Conference on Learning Representations_ , 2025. URL https://openreview.net/forum?id=yAzN4tz7oI.

- [23] Zhen Liu, Xinyu Ning, Zhe Hu, Xinxin Xie, Weize Li, Zhipeng Tang, Chongyu Wang, Zejun Yang, Hanlin Wang, Yitong Liu, and Zhongzhu Pu. Goal2Skill: Long-horizon manipulation with adaptive planning and reflection. _arXiv preprint arXiv:2604.13942_ , 2026.

<!-- End of PDF page 12 -->

- [24] Qian Long, Yueze Wang, Jiaxi Song, Junbo Zhang, Peiyan Li, Wenxuan Wang, Yuqi Wang, Haoyang Li, Shaoxuan Xie, Guocai Yao, Hanbo Zhang, Xinlong Wang, Zhongyuan Wang, Xuguang Lan, Huaping Liu, and Xinghang Li. Scaling world model for hierarchical manipulation policies. _arXiv preprint arXiv:2602.10983_ , 2026.

- [25] Hao Luo, Ye Wang, Wanpeng Zhang, Sipeng Zheng, Ziheng Xi, Chaoyi Xu, Haiweng Xu, Haoqi Yuan, Chi Zhang, Yiqing Wang, Yicheng Feng, and Zongqing Lu. Being-H0.5: Scaling human-centric robot learning for cross-embodiment generalization. _arXiv preprint arXiv:2601.12993_ , 2026.

- [26] Mitsuhiko Nakamoto, Oier Mees, Aviral Kumar, and Sergey Levine. Steering your generalists: Improving robotic foundation models via value guidance. In Pulkit Agrawal, Oliver Kroemer, and Wolfram Burgard, editors, _Proceedings of The 8th Conference on Robot Learning_ , volume 270 of _Proceedings of Machine Learning Research_ , pages 4996–5013. PMLR, 2025. URL https: //proceedings.mlr.press/v270/nakamoto25a.html.

- [27] NVIDIA. Isaac GR00T N1.7: An open reasoning visionlanguage-action model for humanoid robots. https:// huggingface.co/blog/nvidia/gr00t-n1-7, 2026. NVIDIA Isaac GR00T N1.7; foundational report: GR00T N1, arXiv:2503.14734.

- [28] Octo Model Team. Octo: An open-source generalist robot policy, 2024.

- [29] OpenAI. Openai o1 system card. _arXiv preprint arXiv:2412.16720_ , 2024.

- [30] Jeongeun Park, Jihwan Yoon, Byungwoo Jeon, Juhan Park, Jinwoo Shin, Namhoon Cho, Kyungjae Lee, Sangdoo Yun, and Sungjoon Choi. Hierarchical vision language action model using success and failure demonstrations. _arXiv preprint arXiv:2512.03913_ , 2025.

- [31] Physical Intelligence. Knowledge insulating visionlanguage-action models: Train fast, run fast, generalize better, 2025.

- [32] Physical Intelligence. _π_ 0 _._ 5: A vision-language-action model with open-world generalization, 2025.

- [33] Physical Intelligence. _π_ 0 _._ 7: A steerable generalist robotic foundation model with emergent capabilities, 2026.

- [34] Qwen Team. Qwen3.5. https://huggingface.co/ docs/transformers/model <u>doc/qwen3 5,</u> 2026. Qwen3.5 model family.

- [35] Lucy Xiaoyang Shi, Brian Ichter, Michael Equi, Liyiming Ke, Karl Pertsch, Quan Vuong, James Tanner, Anna Walling, Haohuan Wang, Niccolo Fusai, et al. Hi Robot: Open-ended instruction following with hierarchical vision-language-action models. _arXiv preprint arXiv:2502.19417_ , 2025.

- [36] Mustafa Shukor, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans, Steven Palma, Adil Zouitine, Michel Aractingi, Caroline Pascal, Martino Russi, Andres Marafioti, et al. SmolVLA: A vision-languageaction model for affordable and efficient robotics. _arXiv_

_preprint arXiv:2506.01844_ , 2025.

- [37] Ajay Sridhar, Jennifer Pan, Satvik Sharma, and Chelsea Finn. Scaling up memory for robot control via experience retrieval. In _The Fourteenth International Conference on Learning Representations_ , 2026. URL https://openreview.net/forum?id=1dH4ARGdwD.

- [38] Yilin Wu, Thomas Tian, Gokul Swamy, and Andrea Bajcsy. From foresight to forethought: VLM-in-theloop policy steering via latent alignment. In _Proceedings of Robotics: Science and Systems_ , 2025. doi: 10.15607/RSS.2025.XXI.076.

- [39] X-VLA Team. X-VLA: Soft-prompted transformer as scalable cross-embodiment vision-language-action model, 2025.

- [40] Yanting Yang, Shenyuan Gao, Qingwen Bu, Li Chen, and Dimitris N. Metaxas. Seeing farther and smarter: Value-guided multi-path reflection for VLM policy optimization. In _2026 IEEE International Conference on Robotics and Automation (ICRA)_ , 2026.

- [41] Yi Yang, Zhihong Liu, Siqi Kou, Yiyang Chen, Yanzhe Hu, Jianbo Zhou, Boyuan Zhao, Zhijie Wei, Xiao Xia, Xueqi Li, Pengfei Liu, and Zhijie Deng. Worldlanguage-action model for unified world modeling, language reasoning, and action synthesis. _arXiv preprint arXiv:2606.05979_ , 2026.

- [42] Zhilong Zhang, Wenyu Luo, Haonan Wang, Yifei Sheng, Yidi Wang, Hanyuan Guo, Haoxiang Ren, Xinghao Du, Yuhan Che, Tongtong Cao, Lei Yuan, and Yang Yu. Anticipation-VLA: Solving long-horizon embodied tasks via anticipation-based subgoal generation. _arXiv preprint arXiv:2605.01772_ , 2026.

- [43] Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, et al. RT-2: Vision-languageaction models transfer web knowledge to robotic control. In _Conference on Robot Learning_ , pages 2165–2183. PMLR, 2023.

## APPENDIX A AUTHOR CONTRIBUTIONS

**High-Level Policy Training and Research:** Xiaowei Cai, Jingxiao Chen, Xinchen Li, Yifan Li, Yi Liu, Jianlan Luo, Junwen Miao, Ruiqi Ni, Buqing Nie, Jiaxu Wang, Dafeng Wei, Pengwei Xie, Pu Yang, Hangjian Ye, Xiangyu Yue, and Jinyu Zhang.

**Low-Level Policy Training and Research:** Xiaowei Cai, Bingao Chen, Jingxiao Chen, Jingshun Huang, Yi Liu, Jianlan Luo, Junwen Miao, Dafeng Wei, Dongming Wu, Hangjian Ye, Jinyu Zhang, and Pengfei Zhou.

**Training Infra:** Peiqi Wang, Sen Wang, and Qinglin Zhang.

**Robot Infra:** Tengyu Hou, Dong Li, Zhongyuan Liu, and Xiaoyan Wang.

**Writing and Illustration:** Xiaowei Cai, Yunuo Cai, Jingxiao Chen, Zhi Chen, Yi Liu, Jianlan Luo, Junwen Miao, Buqing

<!-- End of PDF page 13 -->

Nie, Dafeng Wei, Dongming Wu, and Jinyu Zhang.

**Data Collection and Deployment:** Zhi Chen, Siyuan Feng, Han Jiang, Runkun Ju, Shaowei Li, Mingjie Pan, Xinlin Ren, Jianheng Song, and Yue Zhou.

**Data Infra:** Mingxiang Li, Xueyong Zhao and Yue Zhou.

**First Authors:** Jinyu Zhang and Yi Liu.

**Corresponding Author:** Jianlan Luo.


![](assets/paper_tau0-vla-2026.pdf-0014-05.png)


We gratefully acknowledge Hongwen Cai, Haoran Chen, Yebin Chen, Zhibo Cui, Zihao Fan, Tao Gao, Yanan Gong, Minghao Gui, Bin He, Xinyan Hou, Lei Hua, Shuaishuai Li, Tong Li, Qingxiang Liu, Xiang L¨u, Zongbao Song, Yuna Sun, Zhe Sun, Zhifeng Tang, Fuming Tian, Jiahao Wan, Hanying Wang, Hao Wang, Sijing Wang, Tao Wang, Le Yuan, Guozhu Zhang, Hao Zhang, Tianbao Zhang, Xinbo Zhi, Mingyan Zhou, and Rongxing Zhu for their valuable contributions to data creation, annotation, and related data operations. We also thank Yifei Chen for assistance with video recording and editing.


![](assets/paper_tau0-vla-2026.pdf-0014-07.png)


This appendix provides additional details on the robot platforms, physical-robot evaluation protocol, low-level training data, and high-level policy data construction.

## _A. Robot Platform Details_

**AGIBOT G1.** AGIBOT G1 is a wheeled humanoid with an omnidirectional four-wheel-steering base, dual 7-DoF arms, and modular end-effectors. Our setup uses parallel-jaw grippers. Its sensing suite includes head-mounted RGB-D and fisheye cameras and a wrist-mounted camera on each arm. **ARX AC One.** ARX AC One is a bimanual platform with dual 6-DoF X5 arms and parallel-jaw grippers. We use wristmounted cameras together with a fixed central camera.

**Franka Research 3.** Our bimanual Franka Research 3 setup comprises two 7-DoF torque-controlled arms with custom 3Dprinted grippers and three RGB cameras. Across all three embodiments, the policy receives multi-view RGB images and proprioceptive state, and the native robot commands are represented in the unified 40-dimensional action layout defined in Section IV-C.

## _B. Unified State and Action Representation_

The canonical state vector **s** _t_ and each action vector **a** _t_ + _j_ , for _j ∈{_ 0 _, . . . , H −_ 1 _}_ , lie in R<sup>40</sup> . The stacked action chunk therefore lies in R<sup>_H×_40</sup> . State and action vectors use the slot ordering shown in Table IV. State channels store the current values, and the end-effector and joint action channels use the relative parameterization defined below. The left and right arm blocks follow each embodiment’s native kinematic joint order. A robot with fewer than eight joints per arm fills the leading entries and masks the remainder.

TABLE IV: **Canonical 40-D state and action layout.** Dimensions are one-indexed. For a rotation matrix **R** = [ **r** 1 _,_ **r** 2 _,_ **r** 3], we use Rot6D( **R** ) = [ **r**<sup>_⊤_</sup> 1<sup>_,_</sup><sup>**r**</sup><sup>_⊤_</sup> 2<sup>]</sup><sup>_⊤_.</sup>

|Coordinates|Dimensions|State representation|
|---|---|---|
|Left EEF position|1–3|Cartesian position in meters<br>|
|Left EEF orientation|4–9|Rot6D(**R**<sup>_L_</sup>)|
|Right EEF position|10–12|Cartesian position in meters<br>|
|Right EEF orientation|13–18|Rot6D(**R**<sup>_R_</sup>)|
|Left gripper|19|native opening coordinate|
|Right gripper|20|native opening coordinate|
|Waist|21–22|two native coordinates|
|Planar base velocity|23–24|two native coordinates|
|Left arm joints|25–32|_q_<sup>_L_</sup><br>1 <sup>_, . . . , qL_</sup><br>8 <sup>in radians</sup><br><br>|
|Right arm joints|33–40|_q_<sup>_R_</sup><br>1 <sup>_, . . . , qR_</sup><br>8 <sup>in radians</sup>|



**Relative action encoding.** Both end-effector and joint actions are expressed relative to the current state. End-effector actions use position and rotation deltas from the current pose, with the position delta expressed in the current end-effector frame. Joint actions use angular offsets from the current joint configuration. **Masks.** Separate state and action masks identify the valid dimensions for each embodiment. They activate the applicable end-effector or joint slots, set unavailable dimensions to zero, and exclude inactive action dimensions from the training loss. **Control metadata.** We serialize _η_ together with the language command _ct_ using the following prompt. The robot type, control mode, and whole-body fields comprise _η_ , while the final field contains _ct_ :

You are controlling a robot. Robot type: _<_ embodiment _>_ Control mode: _<_ eef/eef_wbc/joint _>_ Whole-body control: _<_ enabled/disabled _>_ Task: _<_ language command _>_


![](assets/paper_tau0-vla-2026.pdf-0014-18.png)


The expectation is over training examples, Gaussian noise, and _τ_ = 0 _._ 001 + 0 _._ 999 _x_ with _x ∼_ Beta(1 _._ 5 _,_ 1). The implementation caps each normalized per-sample loss at 100 before batch averaging. At inference, we initialize the chunk from masked Gaussian noise and integrate the learned field from _τ_ = 1 to _τ_ = 0. We apply the action mask to the current noisy chunk before every velocity-field evaluation and to the final chunk after integration. The reported policies use an action horizon of _H_ = 30 and ten uniform Euler updates.

## _C. Physical-Robot Evaluation Protocol_

**Success rate.** We apply the same task-specific terminal conditions and milestone annotations to every method. SR requires every required annotated subtask and the terminal condition

<!-- End of PDF page 14 -->

to be completed. A required subtask completed after one or more failed autonomous attempts still satisfies the SR criterion, provided that no required subtask is skipped and no taskspecific prohibited action occurs.

**Progress.** For each task, we construct a directed acyclic prerequisite graph _G_ = ( _V, E_ ) from its original fine-grained annotations. Each index _i ∈V_ identifies one required milestone. A directed edge ( _i, j_ ) _∈E_ means that milestone _i_ must reach its required task state before milestone _j_ is eligible for credit. Let Anc( _i_ ) denote all direct and indirect prerequisites of milestone _i_ .

For trial _r_ , let _er,i_ = 1 when every milestone in Anc( _i_ ) has reached its required task state before milestone _i_ , and let _er,i_ = 0 otherwise. The local credit _αr,i ∈{_ 0 _,_ 0 _._ 5 _,_ 1 _}_ is 1 for first-attempt completion, 0 _._ 5 for completion after one or more failed autonomous attempts or for a predefined task-specific partial-completion state, and 0 otherwise. The prerequisiteaware credit is


![](assets/paper_tau0-vla-2026.pdf-0015-03.png)


A completed retry can therefore receive half credit while still enabling later milestones. A partial-completion state enables descendants only when it meets the task-specific prerequisite condition. A skipped required milestone receives no credit and blocks only its descendants. Milestones on independent branches remain eligible. For example, failure to add the first topping in Make Milk Tea does not prevent credit for a correctly added second topping. When an execution record groups multiple subtasks, we resolve the outcomes at the original subtask granularity using the accompanying actionlevel annotations.

For a task with _|V|_ annotated milestones and _T_ evaluated trials, progress is


![](assets/paper_tau0-vla-2026.pdf-0015-06.png)


The graph sizes used in Table I are _|V|_ = 25, 14, 22, and 13 for Clean Room, Prepare Ingredients, Tomato and Egg Stir Fry, and Make Milk Tea, respectively. Book Organization uses _|V|_ = 3, and Collect Laundry uses _|V|_ = 5. The three Tidy Makeup Table groups use _|V|_ = 2, 2, and 4 and are evaluated separately.

**Task-specific rules.** Repeated salt addition in Tomato and Egg Stir Fry is a prohibited action rather than an autonomous retry. It invalidates SR and assigns 0 _._ 5 progress credit to the saltaddition subtask. Other autonomously completed subtasks are scored according to the prerequisite graph.

**Trial accounting.** Every method–task entry in Table I uses ten independently collected physical-robot trials.

Each physical-robot trial uses a fixed wall-clock limit measured from the start of autonomous execution. The taskspecific limits are summarized in Table V.

**Reset policy.** Before every trial, the robot is returned to the task-specific standardized initial pose, and all task objects, tools, articulated furniture, and consumables are restored to

TABLE V: **Maximum duration of each physical-robot trial.**

|Task|Maximum duration|
|---|---|
|Clean Room|20 min|
|Prepare Ingredients|20 min|
|Tomato and Egg Stir Fry|20 min|
|Make Milk Tea|10 min|
|Book Organization|5 min|
|Collect Laundry|5 min|
|Tidy Makeup Table (each group)|5 min|



their designated initial states. Object positions are either restored to fixed designated locations or, when a task specifies randomized placement, resampled within the same predefined range for every method. All state carried over from the preceding trial is cleared.

**Termination conditions.** A trial terminates when the task success conditions are satisfied, the task-specific time limit is reached, or the scene enters an irrecoverable state. Only the first case is counted as success. All other termination cases are counted as failures. For progress, only valid milestones completed before termination receive credit, following the scoring rule in the main text.

## _D. Low-Level Data Details_

The low-level training corpus combines approximately 23 _._ 4K hours of internally collected demonstrations with 16 _._ 7K hours from public robot datasets. The internal data comprise 21 _._ 9K hours on AGIBOT G1, 585 hours on AGIBOT G2, 578 hours on ARX AC One, and 347 hours on Franka. The public data include 9 _._ 25K hours of open-source UMI data together with additional selected datasets, supplying further embodiments, tasks, and contact-rich manipulation skills.

We additionally interleave multimodal instruction-following and robot-centric perception examples with the action data. All sources follow a consistent sampling and preprocessing pipeline throughout low-level training; the multimodal samples serve as auxiliary supervision for preserving the backbone’s vision-language capabilities.

## _E. High-Level Policy Data Construction_

Supervision for the high-level policy is generated automatically from task instructions, stage descriptions, executable subtask annotations, segmented demonstrations, and videos. In the data pipeline, we denote these three instruction levels as L1, L2, and L3, respectively. Construction proceeds in three steps. **(1) Pre-labeling.** Using google/Gemma4-31B-it as the annotator, we generate a <think> field that summarizes scene state, progress, constraints, and failures, together with a <memory> field that compresses completed stages while retaining detail for the active subtask. The input includes the carried memory and previously committed subtask. The target <memory> is aligned with the current observation, while the target <subtask> reuses the executable-subtask annotation for what should be executed next. **(2) Keyframe extraction.**

<!-- End of PDF page 15 -->

Following each subtask’s temporal range, we extract three synchronized views (top_head, hand_left, hand_right) with ffmpeg. **(3) VQA assembly.** We combine the synchronized views with the generated fields and executable-subtask annotations to form structured training examples for the highlevel policy.

To make the high-level policy robust to deployment-time memory misalignment, we perturb _only_ the input memory while reading the corrected target from the demonstration. This procedure derives the five instance families in Table VI without additional annotation. Here _Mn_ denotes the memory upon entering segment _n_ . The _within-subtask_ family is the aligned case, and the remaining families target distinct failure modes.

For _error-think_ , the <think> field first flags the failure and then repairs memory according to the failure type. Recoverable failures, such as an empty grasp, are retried with memory unchanged. Failures that undo prior progress, such as a dropped object, trigger rollback to the preceding subtask. Samples marked restorable=false are skipped, and _rollback_ instances are capped at 10–15% to avoid teaching the highlevel policy to distrust correct memory. We additionally apply six-dimensional visually grounded augmentation to the L1/L2 task and stage instructions to improve instruction diversity and zero-shot steerability.

**Quality control.** Quality control operated in two phases. During prompt development, fixed-seed stratified human spotchecks on small batches surfaced failure modes, chiefly crosstask contamination, where the annotator hallucinated action constraints from an unrelated task (e.g. milk-tea constraints appearing under a fruit-shelving task). All visual prompts embed a verification rule that requires omitting any attribute that cannot be confirmed from the images (“when in doubt, leave it out”), and each identified failure mode was hardened into a programmatic filter applied at scale: a cross-task contamination filter, an empty-<subtask> detector, and a memorycontamination filter. Once the prompts and filters were fixed, the full corpus was pre-labeled and filtered programmatically, without further per-sample human inspection.

The contamination filter builds a per-episode noun vocabulary from the L1 and L3 annotations and rejects a constraint whose operated object lies entirely outside this vocabulary, or whose slot nouns overlap it by less than 20%. Tightening the criterion from a naive disjointness test (discard rate 0 _._ 33%, with residual contamination) through an over-aggressive ratio threshold (discard rate 2 _._ 32%, which removed many valid but verbose constraints) to the final two-rule filter with an expanded stop-word list (discard rate 0 _._ 42%) left zero residual contamination on a 3 _._ 26M-sample validation set. Structural dirty-data filtering additionally removes empty-<subtask> episodes and the downstream memory corruption they induce, using a three-tier policy by per-task dirty ratio (tasks above 90% dirty are excluded outright, tasks between 10% and 90% are filtered at the episode level, and tasks below 10% have only their empty-<subtask> samples removed). This discards 11 _._ 74% of episodes and retains 40 _._ 4M (88 _._ 26%) clean

samples. The held-out evaluation set for the high-level policy is drawn from frames unseen in training.

## _F. Asynchronous Dual-System Serving_

The main text describes the logical dependency between high-level subtask generation and low-level action generation as a sequential inference step. In deployment, we pipeline these stages asynchronously because a high-level decision is far slower than one low-level control period. A background worker continuously recomputes the next subtask for each active episode and publishes the generated subtask to a perepisode cache (refreshed about every 1 s). The control loop only reads this cache: each tick takes the cached subtask as the current language command and returns an action chunk immediately, without waiting for the high-level policy. A slow or failed high-level decision thus only delays the next command refresh, while _πθ_ keeps executing the current cached subtask at the control rate ( _∼_ 30 Hz).

## _G. Book Organization Task Settings_

The Book Organization task contains four books of different heights and four fixed slots. At the beginning of each trial, the four books are placed in a shuffled order. At each step, the robot can use its two arms to exchange the positions of any two books. The task is completed when all four books are arranged from tallest to shortest. The high-level policy must therefore predict which pair of books should be exchanged next to make progress toward the target ordering. In the in-domain setting, the initial arrangement of the four books has appeared in the training data. In the out-of-domain (OOD) setting, the initial arrangement of the four books has not appeared in the training data. The OOD setting is designed to evaluate the robustness of the high-level policy to observations arising from initial book arrangements not encountered during training.

## _H. High-Level Evaluation Protocol_

The high-level evaluation isolates decision making from physical execution using annotated samples collected at subtask boundaries. Each sample contains the task instruction, current robot observation, execution memory when applicable, the previously committed subtask, and the ground-truth next subtask. Within each comparison, all methods receive the same samples and produce the same next-subtask output format. The memory ablation changes only whether execution memory is provided. The comparison among Plan Once, Best-of- _N_ , and TTC instead holds the evaluation inputs and output interface fixed while varying the decision-time inference procedure.

## _I. LLM-as-a-Judge Protocol_

For the open-loop TTC evaluation, each sample contains a robot observation from a particular stage of a task and the corresponding ground-truth next subtask. Each method predicts the next subtask from the same input, and an LLM judge evaluates whether the prediction is correct with respect to the ground truth.

<!-- End of PDF page 16 -->

TABLE VI: High-level instance families, synthesized by perturbing only the input memory while reading the corrected target from the demonstration. _Mn_ is the memory upon entering segment _n_ . A single unified <think>/<memory>/<subtask> format instantiates all of them at zero extra annotation.

|Family|Sampling position|Input _→_target memory|Target subtask|Deployment failure countered|Mix|
|---|---|---|---|---|---|
|_within-subtask_|anywhere in seg. _n_|_Mn →Mn_|seg. _n_|— (aligned, normal progression)|58%|
|_transition_|tail of seg. _n_|_Mn →Mn_+1|seg. _n_+1|starting a new subtask after completion|15%|
|_catch-up_|head of seg. _n_|_Mn−_1 _→Mn_|seg. _n_|memory lag (behind the visual state)|10%|
|_rollback_|late in seg. _n_|_Mn_+1_...n_+3 _→Mn_|retry seg. _n_|memory run-ahead (over-optimistic)|12%|
|_error-think_|annotated failure frame|_Mn →_type-dependent|recovery step|unnoticed execution failure|5%|



Exact string matching can substantially underestimate performance because the same executable subtask may be expressed with different wording. We therefore use GPT-5.4 as a semantic judge. For each sample, the judge is given the task goal, the ground-truth next subtask, and the predicted next subtask, and assigns one of three labels: _equivalent_ , _adjacent_ , or _wrong_ .

A prediction is _equivalent_ only if it describes the same immediate physical state transition as the ground truth. The required arm, action, object, destination, and relevant material state must agree, while synonyms and harmless differences in wording or granularity are allowed. A prediction is _adjacent_ if it is a physically valid preceding, following, or reorderable step toward the same task goal, but does not match the immediate transition specified by the ground truth. A prediction is _wrong_ if it uses an incompatible arm, action, object, destination, or state, hallucinates an action, or incorrectly terminates or resets the task. We do not allow the judge to infer omitted objects, destinations, or second-arm actions.

Only _equivalent_ predictions are counted as successful; _adjacent_ is retained as a diagnostic category. Each unique (goal, ground truth, prediction) tuple is evaluated by two separate temperature-zero calls. Disagreements receive a third judgment; if all three labels differ, a fourth judgment is obtained. The final label is determined by majority vote.

prompt You are a strict evaluator of the immediate next executable subtask in a robot plan. Given GOAL, GROUNDTRUTH next subtask, and MODEL-PREDICTED subtask, output one label.

equivalent: the same immediate physical state transition. Required arm(s), action, object(s), destination, and material state must match. Synonyms, capitalization, and harmless wording or granularity differences are allowed.

adjacent: a physically valid preceding, following, or reorderable step for the same goal, but it is a different immediate state transition from the ground truth.

## Classify.

## _J. Adaptive Routing Details_

At deployment, the high-level policy and low-level policy operate in the closed loop summarized in Algorithm 1. At inference step _t_ , the proposal model produces _zt_<sup>dir</sup> and _Mt_ from _ht_ . We compute the routing decision from token logits already produced by this forward pass, so routing requires no additional model invocation.

Let _pi_ be the probability assigned to the generated token at position _i_ . Let _λ_<sup>(1)</sup> _i_ and _λ_<sup>(2)</sup> _i_ be the largest and second-largest logits at that position, and define their margin as


![](assets/paper_tau0-vla-2026.pdf-0017-12.png)


The router uses the mean generated-token probability and the mean logit margin within the <memory> field:


![](assets/paper_tau0-vla-2026.pdf-0017-14.png)


where _I_ all indexes all generated tokens and _I_ mem indexes the tokens in the <memory> field. The binary routing decision is


![](assets/paper_tau0-vla-2026.pdf-0017-16.png)


where **1** [ _·_ ] is the indicator function and _δ_ all and _δ_ mem are routing thresholds. The statistics and routing rule are shared across tasks. The thresholds are calibrated separately for each task on held-out validation data.

When _gt_ = 0, the system sends _zt_<sup>dir</sup> directly to the lowlevel policy. When _gt_ = 1, it searches over candidate subtasks, predicts and scores their visual outcomes, and conditions the reflective model on the retained branches to generate _zt_<sup>_⋆_.</sup>

wrong: incompatible arm, action, object, destination, or state; hallucinated/nonsensical action; or done/reset when the ground truth is another action.

Do not infer missing material objects, destinations, or a missing second-arm action. Respond only as compact JSON:”label”:”equivalent—adjacent—wrong”,”reason”:”¡=16 words”.

GOAL: _{_ task goal _}_ GROUND-TRUTH next subtask: _{_ ground <u>truth</u> _}_ MODEL-PREDICTED subtask: _{_ prediction _}_

<!-- End of PDF page 17 -->
