Uni-LaDiR blog figures

Uni-LaDiR blog figures

All sources now follow arXiv:2609.19878v3, September 29, 2026.

Original vector figures

The PDFs were copied from the public v3 arXiv source. results.json records original filenames and SHA-256 checksums.

  • fig1.pdf: paper Figure 1, overview.
  • fig2.pdf: paper Figure 2, method-overview-v45-reference.pdf. Replaces the older blog diagram.
  • paper-sharing.pdf: paper Figure 3, teacher-sharing study.
  • joint-training.pdf: paper Figure 4, staged versus joint training.
  • objective.pdf: paper Figure 5, encoding and latent-target objectives.
  • diffusion.pdf: paper Figure 6, thought-token generation objectives.

render.py exports these PDFs to SVG and PNG. Raster previews are produced before SVG export because the installed PyMuPDF version can change subsequent rendering state on the page. No measured marks or values are changed in the original figure exports. Paper sharing and objective plots retain the original truncated axes.

Replots and tables

  • main.svg and its mobile version show three named methods per setting from Tables 1–2: Qwen2.5-VL-7B, Mirage, and Uni-LaDiR in both VLM groups; OpenVLA, LaST₀, and Uni-LaDiR on RLBench. OpenVLA is a direct-policy comparator, not a same-backbone ablation. render-results.py produces these charts with fixed 0–100 tick ranges. The VLM comparisons use the same backbone and evaluation examples, with method-specific training data and recipes.
  • sharing-mobile.svg replots v3 Figure 3 values as horizontal bars with axes starting at zero. The desktop post uses the paper’s original paper-sharing.svg.
  • interventions.svg and its mobile version replot Table 7 with the intact chain and three interventions simultaneously visible on zero-based axes. The three ablation summary gains come directly from Section 4.4 and Figures 4–6, rather than being reconstructed from rounded scores.

Reported relative gains use the paper’s unrounded aggregation; plotted scores follow displayed rounding. Sharing studies keep backbone, data, teachers, and thought-token dimensions fixed, but do not match encoder parameter counts.

Run with Python, PyMuPDF, and Matplotlib installed:

python3 assets/blog/uni-ladir/source/render.py
python3 assets/blog/uni-ladir/source/render-results.py
python3 assets/blog/uni-ladir/source/render-fig2-animation.py

The Figure 2 animation uses the user’s original fig2-overview-v45-reference.py, copied unchanged to fig2-animation/fig2/, plus the helper definitions it imports. render-fig2-animation.py executes this copy with a temporary output path, omits the four central gradient arrows, and exports the original artwork to fig2-forward.svg and the inline fig2-art.html. The browser restores those four arrows at their original coordinates as separate SVG paths. Inline vector glyphs preserve sharpness when zooming into a panel. The renderer needs ReportLab, PyMuPDF, Pillow, Matplotlib, and the original macOS Helvetica Neue / Times New Roman fonts.

grounding.html and visuals.js show an overview plus seven stages: encoding, continuation, diffusion, inference, continuation gradients, diffusion gradients, and the shared update. Inactive artwork is gray; an optional enlarged stage view crops the same original vector. Playback starts once on entering the viewport, can be paused or stepped manually, and stops offscreen, in hidden tabs, or at the end. Reduced motion disables autoplay; no-JS retains the complete static diagram and gradients. These are explanatory views of one joint training computation, not sequentially trained systems. Teacher traces are training-only.

Blog explanation figures

decision.html and handoff.html in _includes/blog/uni-ladir/ load editable vector reading aids generated by render-explainers.py. Desktop and mobile SVGs use Helvetica Neue / Helvetica / Arial typography and the original overview’s palette. They compare modality-specific versus shared thought formats, and reconstruction versus continuation supervision. Each teacher step is encoded locally; shared weights do not imply joint access to the other teacher steps. These diagrams contain no empirical values or decoded latent semantics. Mobile variants stack panels and increase label sizes; existing numbered paper figures retain their numbering. Asset URLs carry content hashes for cache invalidation.

toc.html and toc.js provide fixed margin navigation from 1280px, and a native collapsible directory below that width. The directory links to explicit section IDs and tracks the current section. The article restores a 780px reading width despite the site’s theme override. Wide figures retain their own viewport-bounded width, with no overlap with the directory. Dense original plots scroll within their own region on phones; the page itself does not overflow horizontally.

Local Chromium checks at 320, 390, 768, 1280, and 1440px verified all images load, no outer horizontal overflow, six RQ sections, working TLDR and method controls, directory deep links/current-section updates, mobile menu collapse, and no figure/sidebar overlap. Screenshots were reviewed at desktop and phone widths.

The optional top-of-post anonymous Like button uses the community Applause endpoint recommended by the upstream project. It stores shared counts remotely and remembers a vote locally per browser; it is not a verified unique-reader count. Failures are surfaced without recording a local success. Its canonical URL is fixed to the post permalink, so cache-busting query strings do not split counts. The live integration check uses a separate _checks/ URL, never the article counter.

Editorial source notes

The opening motivates a common latent interface for multimodal reasoning, continuation prediction for learning its contents, and diffusion for generating a distribution of next thoughts in that space. The text follows the author-provided opening. Figure 5 is explained as two crossed choices: shared vs separate encoding and continuation vs reconstruction. Its RLBench encoder contrasts (73.0→77.3 under reconstruction; 82.0→87.0 under continuation) follow the original plotted values. Continuation average gains follow the paper’s reported aggregation. Public paper links are version-neutral; exact experimental provenance remains recorded in results.json and source PDF checksums.

Task interpretation and label adaptations

render-task-meaning.py produces desktop/mobile SVGs comparing grasping and color sorting from the same illustrative scene. The image is a conceptual example, not decoded latent contents. The teacher encoder is local to each teacher step; tasks shape it through continuation/final-output gradients. Generated thoughts condition directly on the task input and prefix. The discussion distinguishes this training mechanism from a per-task re-encoding of an identical teacher feature.

render-robot-labels.py expands abbreviated robot-state labels in fig1-blog.pdf and paper-sharing-blog.pdf; the original source PDFs and recorded source checksums remain unchanged. It retains the original figure fonts (Helvetica Neue Bold for the overview and Times New Roman for the chart) and changes no measured values. render.py exports these blog copies when present. The sharing replot labels and conceptual figures also spell out robot state.

RQ2 now asks whether unifying more modalities helps. RQ3.1 holds the objective fixed and compares encoders; RQ3.2 holds the encoder design fixed and compares objectives. Each has its own setup, results, and conclusion. The Platonic comparison is prose rather than a table, separating convergence across independently trained models from deliberately learning a thought space through reasoning supervision.

Sitewide page views use Vercount’s public event endpoint (https://events.vercount.one/api/v2/log). The previous CounterAPI integration remained at zero: the readOnly=false string behaved as read-only, and even requests without that option did not persist in follow-up reads. Vercount was verified using a separate _checks/ URL, including two page loads through the actual site script. Local previews make no counting requests. Production sends one page-view event when visible, keyed to the canonical URL without query strings or fragments; it does not create a persistent visitor ID or report unique visitors. A failed request leaves an unavailable indicator instead of inventing a count.