CITS5553 · Data Science Capstone Project
Companion reading
This pathway builds the lecture in the order in which its questions become answerable: first the multimodal problem, then representation objectives, current systems, embodied action, failure, and predictive world models. “Read” papers carry the argument; “skim” papers supply comparisons; “reference” papers settle a narrower point.
Before multimodality
Starting with language models
-
Language Models are Unsupervised Multitask Learners
read
GPT-2 as the clean statement of the lecture's opening move: the next word is a target the data already contains, and one objective absorbs many tasks.
-
Language Models are Few-Shot Learners
skim
Read for the scale claim only: what happens when self-supervised text prediction meets essentially unlimited training material.
Begin here
The multimodal problem
-
Multimodal Deep Learning
read
Start with reciprocal supervision, shared representations, and missing modalities before foundation-model scale obscures the original problem.
-
Multimodal Machine Learning: A Survey and Taxonomy
read
Use representation, alignment, fusion, translation, and co-learning as the grammar for everything that follows.
Core objectives
How contrastive learning works
-
NYU Deep Learning SP21 · Joint embedding methods, contrastive (15-1)
read
Canziani's lecture notes decompose the family into augmentation, backbone, distance, and anti-collapse losses — the same anatomy as this lecture's four questions. Continue into the regularised methods (15-2) for VICReg, BYOL, and the two collapse types.
-
A comprehensive survey on contrastive learning
read
Use it as a map of views, pairs, negatives, losses, and failure modes; return to the primary papers for mechanisms.
-
Representation Learning with Contrastive Predictive Coding
read
Work through InfoNCE as identification of the matching future latent among sampled alternatives.
-
A Simple Framework for Contrastive Learning of Visual Representations
read
The cleanest controlled account of augmentation, projection heads, temperature, and in-batch comparison.
-
Learning Transferable Visual Models From Natural Language Supervision
read
Follow the move from two views of one image to image–text pairs. The demo scores one row of its matrix live — the one-word rival, the learned temperature, and the left/right question it cannot answer.
-
Barlow Twins
skim
See why negatives are one anti-collapse mechanism rather than the definition of joint-embedding learning.
-
VICReg
read
Separate invariance, variance, and covariance as three geometric pressures.
-
DINOv2
read
Self-supervised features strong enough to reuse without labels — the encoder the live demo interrogates patch by patch.
-
DINOv3
reference
The same recipe scaled to 7B with Gram anchoring for stable dense features; its feature visualisations are the ones shown before the demo.
-
LeJEPA
read
Study the isotropic-Gaussian target and SIGReg as an explicit anti-collapse proposal; keep the geometry claim separate from robotics.
-
VISReg
reference
The supplied 2026 paper: a current refinement that separates scale from distributional shape using sliced-Wasserstein sketches.
-
Joint Embedding vs Reconstruction
read
Use its controlled theory to understand when high-variance irrelevant detail favours latent-space prediction.
-
Learning by Reconstruction Produces Uninformative Features for Perception
reference
A sharper reconstruction critique, bounded by its theory and by the observation that masking changes the problem.
Other modalities
Audio and speech
-
wav2vec 2.0
skim
A masked, contrastive identification problem inside one audio stream.
-
Robust Speech Recognition via Large-Scale Weak Supervision
read
Whisper is the corrective case: sequence-to-sequence learning at scale, not contrastive pretraining.
-
CLAP
read
Compare audio–language alignment with both temporal audio prediction and transcription.
-
ImageBind
read
Ask what is gained, and assumed, when image-paired data become the hub for six modalities.
-
Qwen3.5-Omni Technical Report
read
Trace Thinker–Talker, chunk-wise streaming, audio–video timestamps, causal speech rendering, and text–speech token-rate alignment.
Current systems
Vision–language models
-
An Image is Worth 16×16 Words
skim
ViT establishes the move the VLM chapter depends on: an image can enter a transformer as a sequence of patch embeddings.
-
Visual Instruction Tuning
read
LLaVA is the lecture's worked example of how image patches become tokens a language model reads: a projector maps CLIP patch features into the embedding space, trained by caption alignment and then visual instruction data.
-
Flamingo
skim
Perceiver resampling and gated cross-attention into a frozen language model.
-
BLIP-2
read
The Q-Former makes the bottleneck between frozen vision and language models explicit.
-
Qwen3-VL Technical Report
read
Update the architecture story through SigLIP-2, visual-token merging, DeepStack, multidimensional position, and explicit video timestamps.
-
SmolVLM
read
The pretrained VLM (SigLIP encoder, SmolLM2 decoder) that the lecture's VLA starts from — read it before SmolVLA to see exactly what is taken frozen when a VLM is turned into a VLA and fine-tuned on LIBERO.
-
VL-JEPA
read
A controlled alternative to token-space prediction: continuous semantic targets and selective decoding, with explicit limits.
-
Vision-Guided Iterative Refinement for Frontend Code Generation
skim
A practical digital action loop: generate code, render it, inspect the result with a VLM critic, and revise. Read the evaluation and latency limits alongside the improvement.
Embodied action
Vision–language–action models
-
PaLM-E
read
A bridge from multimodal language modelling to continuous sensor and state inputs; do not confuse embodied reasoning with low-level control.
-
RT-1
skim
Establish the large-scale behaviour-cloning substrate before the VLA label.
-
RT-2
read
Understand action-as-token co-fine-tuning and distinguish semantic transfer from motor reliability.
-
Open X-Embodiment
skim
The data substrate: scale through heterogeneous robots, plus normalisation and embodiment mismatch.
-
Octo
skim
A useful non-VLM generalist-policy comparison: which gains come from data and policy pretraining?
-
OpenVLA
read
Fused semantic and local visual features, a language backbone, action tokens, and a practical open fine-tuning route. The blue-cup / pink-cup and reattempt clips on the VLA chapter slides are OpenVLA.
-
π₀
read
Contrast discrete action tokens with a continuous action expert trained by flow matching — the mechanism SmolVLA inherits, and the folding footage in the deck.
-
SmolVLA
read
The model the demo runs and fine-tunes: the frozen SmolVLM above with a ~100M flow-matching action expert — the architecture behind every training tier shown in the lecture.
Evaluation
VLA evaluation and failure modes
-
Evaluating Real-World Robot Manipulation Policies in Simulation
read
Learn what useful ranking correspondence can establish without treating simulation as physical reality.
-
From Intention to Execution
read
The central criticism: correct semantic intention can survive while physical execution fails. Read the benchmark limitations just as closely.
World models
Predictive world models
-
Introduction to Latent Variable Energy-Based Models
read
Read the programme as a diagnosis and proposal, not as experimental proof that a full architecture already works.
-
I-JEPA
read
The image-level proof of concept for target prediction in representation space.
-
V-JEPA
read
Extend feature prediction through time while keeping representation evidence separate from physical understanding.
-
Audio-JEPA
skim
Test whether the JEPA construction is modality-general by comparing it with wav2vec 2.0, CLAP, and Whisper.
-
LeWorldModel
read
Action-conditioned future-latent prediction with explicit Gaussian regulation; track short horizons and offline coverage.
-
When Does LeJEPA Learn a World Model?
reference
The strongest theoretical claim, and a reminder to inspect Gaussian, dimensionality, optimum, and finite-sample assumptions.
-
V-JEPA 2
read
Direct robotics evidence for image-goal planning; the demo uses its action-conditioned energy to score nine named commands and a 343-candidate grid, with the reversed-clip control. Read camera sensitivity, search cost, latency, and horizon limits alongside success.
-
Hierarchical Planning with Latent World Models
read
A current attack on planning horizon through latent subgoals, still bounded by visual goals and selected manipulation settings.