Prosody LabsLectures

CITS5553 · Data Science Capstone Project

Companion reading

This pathway builds the lecture in the order in which its questions become answerable: first the multimodal problem, then representation objectives, current systems, embodied action, failure, and predictive world models. “Read” papers carry the argument; “skim” papers supply comparisons; “reference” papers settle a narrower point.

00

Before multimodality

Starting with language models

  1. Language Models are Unsupervised Multitask Learners read

    GPT-2 as the clean statement of the lecture's opening move: the next word is a target the data already contains, and one objective absorbs many tasks.

  2. Language Models are Few-Shot Learners skim

    Read for the scale claim only: what happens when self-supervised text prediction meets essentially unlimited training material.

01

Begin here

The multimodal problem

  1. Multimodal Deep Learning read

    Start with reciprocal supervision, shared representations, and missing modalities before foundation-model scale obscures the original problem.

  2. Multimodal Machine Learning: A Survey and Taxonomy read

    Use representation, alignment, fusion, translation, and co-learning as the grammar for everything that follows.

02

Core objectives

How contrastive learning works

  1. NYU Deep Learning SP21 · Joint embedding methods, contrastive (15-1) read

    Canziani's lecture notes decompose the family into augmentation, backbone, distance, and anti-collapse losses — the same anatomy as this lecture's four questions. Continue into the regularised methods (15-2) for VICReg, BYOL, and the two collapse types.

  2. A comprehensive survey on contrastive learning read

    Use it as a map of views, pairs, negatives, losses, and failure modes; return to the primary papers for mechanisms.

  3. Representation Learning with Contrastive Predictive Coding read

    Work through InfoNCE as identification of the matching future latent among sampled alternatives.

  4. A Simple Framework for Contrastive Learning of Visual Representations read

    The cleanest controlled account of augmentation, projection heads, temperature, and in-batch comparison.

  5. Learning Transferable Visual Models From Natural Language Supervision read

    Follow the move from two views of one image to image–text pairs. The demo scores one row of its matrix live — the one-word rival, the learned temperature, and the left/right question it cannot answer.

  6. Barlow Twins skim

    See why negatives are one anti-collapse mechanism rather than the definition of joint-embedding learning.

  7. VICReg read

    Separate invariance, variance, and covariance as three geometric pressures.

  8. DINOv2 read

    Self-supervised features strong enough to reuse without labels — the encoder the live demo interrogates patch by patch.

  9. DINOv3 reference

    The same recipe scaled to 7B with Gram anchoring for stable dense features; its feature visualisations are the ones shown before the demo.

  10. LeJEPA read

    Study the isotropic-Gaussian target and SIGReg as an explicit anti-collapse proposal; keep the geometry claim separate from robotics.

  11. VISReg reference

    The supplied 2026 paper: a current refinement that separates scale from distributional shape using sliced-Wasserstein sketches.

  12. Joint Embedding vs Reconstruction read

    Use its controlled theory to understand when high-variance irrelevant detail favours latent-space prediction.

  13. Learning by Reconstruction Produces Uninformative Features for Perception reference

    A sharper reconstruction critique, bounded by its theory and by the observation that masking changes the problem.

03

Other modalities

Audio and speech

  1. wav2vec 2.0 skim

    A masked, contrastive identification problem inside one audio stream.

  2. Robust Speech Recognition via Large-Scale Weak Supervision read

    Whisper is the corrective case: sequence-to-sequence learning at scale, not contrastive pretraining.

  3. CLAP read

    Compare audio–language alignment with both temporal audio prediction and transcription.

  4. ImageBind read

    Ask what is gained, and assumed, when image-paired data become the hub for six modalities.

  5. Qwen3.5-Omni Technical Report read

    Trace Thinker–Talker, chunk-wise streaming, audio–video timestamps, causal speech rendering, and text–speech token-rate alignment.

04

Current systems

Vision–language models

  1. An Image is Worth 16×16 Words skim

    ViT establishes the move the VLM chapter depends on: an image can enter a transformer as a sequence of patch embeddings.

  2. Visual Instruction Tuning read

    LLaVA is the lecture's worked example of how image patches become tokens a language model reads: a projector maps CLIP patch features into the embedding space, trained by caption alignment and then visual instruction data.

  3. Flamingo skim

    Perceiver resampling and gated cross-attention into a frozen language model.

  4. BLIP-2 read

    The Q-Former makes the bottleneck between frozen vision and language models explicit.

  5. Qwen3-VL Technical Report read

    Update the architecture story through SigLIP-2, visual-token merging, DeepStack, multidimensional position, and explicit video timestamps.

  6. SmolVLM read

    The pretrained VLM (SigLIP encoder, SmolLM2 decoder) that the lecture's VLA starts from — read it before SmolVLA to see exactly what is taken frozen when a VLM is turned into a VLA and fine-tuned on LIBERO.

  7. VL-JEPA read

    A controlled alternative to token-space prediction: continuous semantic targets and selective decoding, with explicit limits.

  8. Vision-Guided Iterative Refinement for Frontend Code Generation skim

    A practical digital action loop: generate code, render it, inspect the result with a VLM critic, and revise. Read the evaluation and latency limits alongside the improvement.

05

Embodied action

Vision–language–action models

  1. PaLM-E read

    A bridge from multimodal language modelling to continuous sensor and state inputs; do not confuse embodied reasoning with low-level control.

  2. RT-1 skim

    Establish the large-scale behaviour-cloning substrate before the VLA label.

  3. RT-2 read

    Understand action-as-token co-fine-tuning and distinguish semantic transfer from motor reliability.

  4. Open X-Embodiment skim

    The data substrate: scale through heterogeneous robots, plus normalisation and embodiment mismatch.

  5. Octo skim

    A useful non-VLM generalist-policy comparison: which gains come from data and policy pretraining?

  6. OpenVLA read

    Fused semantic and local visual features, a language backbone, action tokens, and a practical open fine-tuning route. The blue-cup / pink-cup and reattempt clips on the VLA chapter slides are OpenVLA.

  7. π₀ read

    Contrast discrete action tokens with a continuous action expert trained by flow matching — the mechanism SmolVLA inherits, and the folding footage in the deck.

  8. SmolVLA read

    The model the demo runs and fine-tunes: the frozen SmolVLM above with a ~100M flow-matching action expert — the architecture behind every training tier shown in the lecture.

06

Evaluation

VLA evaluation and failure modes

  1. Evaluating Real-World Robot Manipulation Policies in Simulation read

    Learn what useful ranking correspondence can establish without treating simulation as physical reality.

  2. From Intention to Execution read

    The central criticism: correct semantic intention can survive while physical execution fails. Read the benchmark limitations just as closely.

07

World models

Predictive world models

  1. Introduction to Latent Variable Energy-Based Models read

    Read the programme as a diagnosis and proposal, not as experimental proof that a full architecture already works.

  2. I-JEPA read

    The image-level proof of concept for target prediction in representation space.

  3. V-JEPA read

    Extend feature prediction through time while keeping representation evidence separate from physical understanding.

  4. Audio-JEPA skim

    Test whether the JEPA construction is modality-general by comparing it with wav2vec 2.0, CLAP, and Whisper.

  5. LeWorldModel read

    Action-conditioned future-latent prediction with explicit Gaussian regulation; track short horizons and offline coverage.

  6. When Does LeJEPA Learn a World Model? reference

    The strongest theoretical claim, and a reminder to inspect Gaussian, dimensionality, optimum, and finite-sample assumptions.

  7. V-JEPA 2 read

    Direct robotics evidence for image-goal planning; the demo uses its action-conditioned energy to score nine named commands and a 343-candidate grid, with the reversed-clip control. Read camera sensitivity, search cost, latency, and horizon limits alongside success.

  8. Hierarchical Planning with Latent World Models read

    A current attack on planning horizon through latent subgoals, still bounded by visual goals and selected manipulation settings.