AI

Physical Intelligence: How VLA Models Are Giving Robots a Human Sense of Sight

TL;DR A rule of thumb emerging from physical AI research: 50-200 human demonstrations, collected via teleoperation or kinesthetic teaching (physically moving the robot's arm), is typically sufficient to train a VLA model to perform a new manipulation task reliably. This is orders of magnitude fewer than classical supervised learning requires — because VLA models bring vast world knowledge from pretraining.

Read this article as text (accessible version)
// Physical Intelligence · Edge Robotics · 2026 · 3,900 words · 4 interactive labs · April 2026

Physical Intelligence:
How VLA Models Are
Giving Robots a
Human Sense of Sight

For sixty years, robotics meant rigid programming. You told a robot arm exactly which joint to move how many degrees, at what speed, in what sequence. If someone put the parts bin 10 centimeters to the left of where the program expected, the robot either crashed or stopped. That era is ending. VLA models change everything — and the implications ripple through every industry that touches the physical world.

Read the Deep Dive ↓ Open Robotics Lab 🤖 Vision → Language → Action Imitation Learning Edge AI Chips · µs Inference LBMs replace rigid programming Smart agriculture · maritime · manufacturing // Table of Contents
  1. Demystifying VLA Models
  2. Imitation Learning: Watch & Do
  3. The Edge Computing Revolution
  4. Commercial Applications Now
  5. How It All Connects
  6. FAQ

A researcher at a Japanese robotics lab sat down in front of a robotic arm and folded a towel — once. Just once, in full view of the robot's cameras. Then she stepped back and watched the robot fold the next towel. And the next. And forty-seven towels after that, all slightly different sizes, colors, and starting positions. The robot had never been programmed with a single line of code specifying what "folding a towel" means. It learned by watching.

This is what Physical Intelligence looks like in practice — and it's not just a lab demonstration anymore. The underlying architecture, Vision-Language-Action models, is making its way into commercial production systems across agriculture, maritime maintenance, and precision manufacturing. Understanding how this technology works — and where it's heading — is no longer optional for engineers who want to stay relevant in the next decade.

// 01Demystifying VLA Models: Vision + Language + Action

The name is the architecture: Vision-Language-Action models take three distinct input and output modalities and bind them into a unified system. Vision: continuous camera feeds providing spatial awareness of the environment — object positions, orientations, occlusions, lighting conditions. Language: natural language task descriptions, either from a human operator ("pick up the red cylinder and place it in the left bin") or from a higher-level planning system. Action: the output — not text, not images, but precise motor commands: joint angles, gripper forces, velocity vectors, trajectory waypoints.

What makes VLA models architecturally distinct from standard LLMs is the action head. A language model's output space is a probability distribution over tokens — words. A VLA model's output space includes a policy head that generates continuous motor control signals. The model must simultaneously understand semantic meaning (what "the red cylinder" refers to in the scene) and physical affordances (how to grasp an object without dropping it based on its shape, material, and current orientation). This requires models that can reason about physics, geometry, and causality — not just language patterns.

The most widely studied VLA architectures — Google DeepMind's RT-2, Physical Intelligence's π0 (pi-zero), and OpenVLA — all share a similar structural approach: a pre-trained vision-language backbone (typically a large vision-language model like PaLI or LLaMA) fine-tuned or adapted with an action decoder that maps vision-language representations to motor outputs. This transfer of knowledge from language-vision pretraining to robot control is the key insight — the model's understanding of "pick up," "place," "sort," and "fold" comes from training on internet-scale text and image data, then it's adapted to specific robotic embodiments with comparatively small amounts of robot-specific training data.

Here's the thing most robotics tutorials miss about VLA models: the language component isn't just for taking commands — it's a semantic backbone for generalization. When a robot trained on VLA understands the semantic concept of "fragile" from language pretraining, it can apply that understanding to novel objects it's never handled before. A robot sees a glass sculpture and, without specific training on glass sculptures, understands to reduce grip force and move slowly. The language model's world knowledge becomes physical intuition. This cross-domain knowledge transfer is what makes VLA fundamentally different from classical robotic programming.

💡 The Tokenization Trick Behind Action Outputs

One elegant architectural detail: many VLA models convert continuous motor actions into discrete tokens using a process called action tokenization. Joint angles (continuous values like 34.7°) are binned into 256 possible discrete values and treated as vocabulary tokens — just like words. This means the model generates actions using the same transformer decoder mechanism it uses for generating text. The beauty: you can train a single model end-to-end for both language generation AND motor control using the same loss function and optimizer. The downside: discretization introduces quantization error, which is why high-precision industrial tasks may still require hybrid architectures with a continuous control head for the final refinement pass.

vla_inference.py — querying a VLA model for motor commands
import torch
from transformers import AutoProcessor, AutoModelForVision2Seq
from PIL import Image
import numpy as np

# Load OpenVLA (open-source VLA model)
processor = AutoProcessor.from_pretrained("openvla/openvla-7b", trust_remote_code=True)
vla = AutoModelForVision2Seq.from_pretrained("openvla/openvla-7b",
 attn_implementation="flash_attention_2",
 torch_dtype=torch.bfloat16).to("cuda")

# Input: camera frame + natural language instruction
image = Image.open("robot_camera_frame.jpg")
instruction = "Pick up the red cylindrical object and place it into the left bin"

inputs = processor.apply_chat_template(
 [{"role": "user",
 "content": [{"type": "image"},
 {"type": "text", "text": instruction}]}],
 add_generation_prompt=True, tokenize=True, return_tensors="pt"
).to("cuda")

# Output: 7-DoF robot action (x, y, z, roll, pitch, yaw, gripper)
action = vla.predict_action(**inputs, unnorm_key="bridge_orig")
# Returns numpy array: [dx, dy, dz, droll, dpitch, dyaw, gripper_open]
# Example: [-0.012, 0.034, -0.008, 0.002, -0.001, 0.007, 0.92]
# → Move slightly left (+x), forward (y), down (z), open gripper 92%

print(f"Action: {action}")
robot.execute_action(action) # Send to robot controller
LLM next token prediction probability distribution diagram showing vocabulary tokens with probability bars and sampling mechanism

// 02From Programming to Imitation Learning: Watch Once, Do Forever

Classical robotic programming is fundamentally a translation problem: a human engineer observes a task, mentally decomposes it into discrete mechanical steps, and translates each step into robot controller instructions — joint torques, coordinate frames, velocity profiles, collision avoidance constraints. This translation is expensive (skilled engineers), slow (months per task), brittle (any environmental change requires reprogramming), and doesn't generalize (a robot programmed to pick 50mm cylindrical caps can't pick 52mm caps without modification).

Imitation learning — also called Learning from Demonstration (LfD) — inverts this process. A human performs the task while the robot watches and records. The robot's model learns to map from observations (what the camera sees) to actions (what the human's hands did) directly, without ever requiring a human to explicitly encode "step 3: move joint 4 to 34.7 degrees." Modern implementations go further: with Action Chunking with Transformers (ACT), robots learn from as few as 50 human demonstrations per task and achieve reliable performance. A task that took months to program classically can now be taught in hours of human teleoperation.

Synthetic simulation takes this even further. Tools like NVIDIA Isaac Sim and Google's MimicGen can take a handful of real human demonstrations and automatically generate thousands of variations: different object positions, lighting conditions, environmental clutter, gripper starting positions. This synthetic augmentation creates training datasets that would be impossible to collect in the real world — robots learning to handle edge cases that humans might only encounter once a month in a production environment. The result is robots that are genuinely robust to the variability of real-world conditions, not brittle programs that fail when someone moves a bin five centimeters.

✅ The Magic Number: 50-200 Demonstrations

A rule of thumb emerging from physical AI research: 50-200 human demonstrations, collected via teleoperation or kinesthetic teaching (physically moving the robot's arm), is typically sufficient to train a VLA model to perform a new manipulation task reliably. This is orders of magnitude fewer than classical supervised learning requires — because VLA models bring vast world knowledge from pretraining. The practical implication: any organization with a few hours of human operator time and a robot arm can now "program" new tasks without a robotics engineer. This democratization is the most underappreciated consequence of VLA models for industrial applications.

collect_demonstrations.py — teleoperation data collection
import numpy as np
from dataclasses import dataclass, field
from typing import List
import h5py, cv2, time

@dataclass
class DemoStep:
 timestamp: float
 rgb_frame: np.ndarray # (480, 640, 3) camera image
 joint_positions: np.ndarray # 7-DoF joint angles (radians)
 gripper_state: float # 0.0 (closed) to 1.0 (open)
 action: np.ndarray # delta joint velocities this timestep
 task_instruction: str

class DemonstrationCollector:
 def __init__(self, task_name: str, save_path: str):
 self.task_name = task_name
 self.save_path = save_path
 self.episodes: List[List[DemoStep]] = []
 
 def record_episode(self, robot, camera, instruction: str):
 episode = []
 print(f"Recording: '{instruction}' - Press Enter when done")
 while not human_says_done():
 step = DemoStep(
 timestamp=time.time(),
 rgb_frame=camera.capture(), # 30 Hz camera
 joint_positions=robot.get_joints(), # current state
 gripper_state=robot.gripper_pos(),
 action=robot.last_teleop_action(), # what human did
 task_instruction=instruction
 )
 episode.append(step)
 time.sleep(1/30) # 30 Hz collection rate
 self.episodes.append(episode)
 print(ff"Recorded {len(episode)} steps. Total: {len(self.episodes)} demos")
 
 def save_dataset(self):
 # Save in LeRobot HDF5 format for VLA fine-tuning
 with h5py.File(self.save_path, 'w') as f:
 for i, ep in enumerate(self.episodes):
 grp = f.create_group(ff'episode_{i}')
 grp.create_dataset('images',
 data=np.stack([s.rgb_frame for s in ep]))
 grp.create_dataset('actions',
 data=np.stack([s.action for s in ep]))

// 03The Edge Computing Revolution: Why the Cloud Can't Run a Robot

Picture this: a robotic surgical assistant pauses for 200 milliseconds in the middle of a delicate procedure because its inference request to a cloud server is waiting for a network response. This scenario isn't hypothetical — it's the fundamental reason why physically-intelligent systems cannot rely on cloud computing. The physical world operates in real time and doesn't wait for network latency to resolve. When a robot arm is moving at operating speed and needs to avoid an unexpected obstacle, it needs to update its control signals in under 10 milliseconds. Cloud inference latency is measured in hundreds of milliseconds. The math doesn't work.

Edge AI hardware is the solution: specialized chips that run neural network inference locally, on the device, without any network dependency. The current leading options each represent a different optimization philosophy. NVIDIA Jetson Orin provides the familiar CUDA programming model in a small form factor — 275 TOPS (Trillion Operations Per Second) at under 60W. Google Coral Edge TPU provides extremely low power consumption for smaller models. Qualcomm AI 100 Ultra combines CPU, GPU, and dedicated AI accelerators for multi-modal inference. And the newest category — neural processing units (NPUs) built directly into microcontrollers — brings AI inference to devices with milliwatt power budgets.

The counterintuitive insight about edge AI for robotics: you don't run the full VLA model on the edge. The architecture is typically hierarchical. A small, fast model runs on the edge chip at high frequency (30-100 Hz) to handle immediate reactive control — obstacle avoidance, real-time force feedback, immediate safety responses. A medium model runs at 10-30 Hz for local task planning — "I've grasped the object; now move it to the target location." And a larger model handles higher-level task understanding less frequently — "What is my current task? What's the next goal state?" The biggest models can offload to local (on-premises) compute when available; they don't need to hit the internet.

⚡ Model Quantization: Fitting a VLA on an Edge Chip

A 7-billion parameter VLA model in float32 requires ~28GB of memory — far beyond any current edge chip. The solution is quantization: converting model weights from 32-bit floats to 8-bit integers (INT8) or 4-bit (INT4), reducing memory footprint by 4-8× with modest accuracy loss. A 7B parameter model quantized to INT4 requires ~3.5GB and runs on an NVIDIA Jetson AGX Orin's 32GB shared memory. This is the practical pipeline: train the full model in bfloat16 on data center GPUs, quantize and distill for edge deployment, and validate that task performance meets acceptable thresholds. Tools like llama.cpp, ExLlamaV2, and NVIDIA TensorRT-LLM handle the quantization pipeline for deployment.

edge_deploy.py — quantized VLA for Jetson Orin
import tensorrt as trt
import pycuda.driver as cuda
import numpy as np, time

# Step 1: Export VLA model to TensorRT for Jetson optimization
# (Run once on development machine, deploy .engine file)
def build_engine(onnx_path: str, engine_path: str, precision: str = "int8"):
 builder = trt.Builder(trt.Logger(trt.Logger.WARNING))
 network = builder.create_network()
 parser = trt.OnnxParser(network, builder.logger)
 
 config = builder.create_builder_config()
 config.set_memory_pool_limit(trt.MemoryPoolType.WORKSPACE, 4 << 30) # 4GB
 
 if precision == "int8":
 config.set_flag(trt.BuilderFlag.INT8) # 4× memory savings
 config.set_flag(trt.BuilderFlag.FP16) # fallback for unsupported ops
 
 # Build and serialize engine (takes ~10 min, runs in ms afterward)
 engine = builder.build_serialized_network(network, config)
 with open(engine_path, 'wb') as f:
 f.write(engine)

# Step 2: Real-time inference loop on Jetson (runs in <10ms per step)
class EdgeVLARunner:
 def __init__(self, engine_path: str):
 with open(engine_path, 'rb') as f:
 self.engine = trt.Runtime(trt.Logger()).deserialize_cuda_engine(f.read())
 self.context = self.engine.create_execution_context()

 def infer(self, image: np.ndarray, instruction_tokens: np.ndarray) -> np.ndarray:
 start = time.perf_counter()
 # Upload inputs, run inference, download action output
 action = self._run_engine(image, instruction_tokens)
 latency_ms = (time.perf_counter() - start) * 1000
 assert latency_ms < 10, ff"Latency {latency_ms:.1f}ms exceeds 10ms budget!"
 return action # 7-DoF motor command, generated offline on-device

// 04Commercial Applications: Beyond Humanoid Hype

The media coverage of physical AI has been heavily skewed toward humanoid robots — Boston Dynamics Atlas, Figure 01, Tesla Optimus. These are technically impressive, but they're not where physical AI is generating commercial value right now. The highest ROI deployments are in three domains that are less photogenic but far more economically important: precision agriculture, maritime vessel maintenance, and adaptive manufacturing.

Precision agriculture is perhaps the most compelling near-term application. Agricultural robots equipped with VLA models can identify individual diseased plant specimens in a field, navigate to them, and administer targeted treatment — without prior programming for every possible plant variety and disease state. The robot learns from agronomist demonstrations rather than requiring an exhaustive rule-based identification system. Harvest CROO Robotics and Carbon Robotics are deploying systems that handle strawberry picking and laser-based weed elimination with human-competitive dexterity and reliability. The economic driver: labor costs in agriculture are rising while crop margins are under pressure. A robot that can be taught a new task by a farm worker in an afternoon, rather than reprogrammed by an engineer over weeks, changes the ROI calculation entirely.

Maritime vessel maintenance is a domain most people don't think about but represents one of the highest-value applications. Ships operating globally can't always reach a drydock when a maintenance issue arises. VLA-equipped drone systems are being deployed that can inspect hull integrity, conduct underwater welding repairs, and perform equipment diagnostics on vessels underway — all from demonstrations given by human divers and maintenance engineers. The latency requirements make this a perfect edge AI use case: offshore vessels have unreliable satellite connectivity, so all AI inference must run on the vessel's own hardware. Adaptive manufacturing closes the loop: traditional robotic assembly lines require weeks of engineering when a new product is introduced. Lines equipped with VLA models can be retrained on a new product variant in hours, by having a skilled assembly worker demonstrate the process.

🔬 The Dexterous Manipulation Gap — Still Unsolved

Physical AI in 2026 has a well-documented capability gap: dexterous manipulation of deformable objects. Folding clothes, handling flexible cables, managing loose biological materials — tasks that require continuous real-time tactile feedback and complex multi-finger coordination — remain significantly harder than rigid object manipulation. Current VLA models achieve strong performance on rigid body tasks but degrade on high-dexterity deformable manipulation. The research frontier is multimodal sensory integration: adding tactile sensor data (from fingertip pressure sensors) and proprioceptive feedback alongside vision to create a richer sensory model. Systems from labs like CMU and UC Berkeley are showing early results, but commercial deployment of dexterous manipulation at scale remains 2-4 years out.


// synthesisHow It All Connects: The Physical AI Stack

VLA models provide the cognitive architecture — the ability to interpret visual scenes, understand natural language instructions, and generate appropriate motor commands. Imitation learning provides the data acquisition mechanism — enabling rapid task acquisition from human demonstrations without classical programming. Edge AI hardware provides the deployment platform — making sub-10ms inference possible on devices without network connectivity. And the commercial applications provide the economic context — the domains where these capabilities create enough value to justify deployment today.

What this means for software development as a discipline is genuinely profound. Writing code that controls physical systems through sensor data and learned policies is a fundamentally different paradigm from writing code that processes data on a server. The abstractions change: instead of function calls and data structures, you're working with perception pipelines, policy networks, safety constraints, and real-time control loops. Engineers who understand both the software stack (model training, quantization, deployment) and the physical domain (kinematics, force control, sensor calibration) will be among the most valuable technical professionals in the next decade.


// FAQFrequently Asked Questions

What is physical intelligence in AI and robotics? + Physical intelligence refers to AI systems that understand and interact with the physical world — not just text or images, but real physical objects, spatial relationships, forces, and motor control. The key capability is the ability to perceive a physical environment (via cameras and sensors), reason about it (what objects are present, what their properties are, what actions are possible), and take physical actions (moving robot joints, applying force, navigating space). Physical intelligence is enabled by Large Behavior Models (LBMs) and Vision-Language-Action (VLA) models that combine visual perception, language understanding, and motor control in a single architecture trained on demonstration data rather than manually programmed rules. What is a VLA model and how does it differ from an LLM? + A VLA (Vision-Language-Action) model takes camera images (Vision) and natural language instructions (Language) as input and outputs robotic motor commands (Action) — joint angles, gripper forces, velocity trajectories. A standard LLM only takes text and outputs text. The critical architectural difference is the action head: VLA models include a policy decoder that generates continuous or discretized motor control signals rather than text tokens. VLA models inherit world knowledge and semantic understanding from pretraining on large-scale vision-language data, then add robotic action capabilities through fine-tuning on demonstration datasets. This transfer learning approach dramatically reduces the amount of robot-specific training data needed for new tasks. What is imitation learning in robotics? + Imitation learning (also called Learning from Demonstration or LfD) is a paradigm where robots learn tasks by observing human demonstrations rather than being manually programmed. A human performs the task (via teleoperation or directly moving the robot's arm) while the robot's sensors record observations and actions. A machine learning model then learns to map from observations (camera frames, sensor data) to actions (motor commands), effectively learning to imitate the human's behavior. Modern approaches like Action Chunking with Transformers (ACT) and diffusion policies can achieve reliable performance from as few as 50-200 demonstrations per task — orders of magnitude fewer than classical data-hungry ML approaches because VLA models bring pre-existing world knowledge from pretraining. Why do robots need edge AI instead of cloud computing? + Physical robotic systems require control loop update rates of 30-1000 Hz with latency under 10ms per inference cycle. Cloud computing adds network round-trip latency of 50-300ms (under ideal conditions) — making it physically impossible to maintain real-time robot control through a cloud connection. Additionally: (1) Network reliability — robots in remote or industrial environments often have unreliable or absent connectivity. (2) Safety — if network connectivity is lost, a cloud-dependent robot must stop; an edge-capable robot can continue operating. (3) Bandwidth — high-frequency camera streams plus 30Hz action outputs generate data volumes that exceed practical wireless bandwidth. Edge AI chips like NVIDIA Jetson Orin run quantized VLA models at <10ms latency locally, solving all three constraints. What industries are already using VLA models and physical AI? + Commercial deployments in 2026 span several sectors. Precision agriculture: robots that identify and treat diseased plants, harvest delicate produce, and navigate unstructured farm terrain — taught by agronomist demonstrations rather than programmed. Warehouse logistics: Amazon, Ocado, and several robotics startups deploy VLA-based systems for item picking and sorting across diverse product catalogues without task-specific programming for each SKU. Maritime maintenance: drone systems for hull inspection and repair on operating vessels without drydock. Semiconductor manufacturing: adaptive assembly robots that handle product changeovers in hours instead of weeks. Restaurant food preparation: several quick service restaurant chains are piloting VLA-based cooking assistants for standardized but variable preparation tasks. What hardware is needed to run VLA models on edge devices? + Current production edge AI hardware for robotics: NVIDIA Jetson AGX Orin (275 TOPS, 64GB memory, ~$800) is the workhorse for most commercial deployments — sufficient for running quantized 7B parameter VLA models in real time. Jetson Orin Nano (40 TOPS, ~$250) suits lower-power/cost applications with smaller models. For ultra-low-latency requirements, NVIDIA Jetson Thor (2,000 TOPS, due 2026) will enable full-precision larger models on-device. Google Coral with Jetson provides specialized acceleration for certain model architectures. The practical pipeline: develop on a data center GPU cluster, quantize to INT8 or INT4, profile on the target Jetson, adjust model size/quantization to meet latency budget, then deploy. How long does it take to train a robot with imitation learning? + With modern VLA fine-tuning approaches: collecting 50-100 human demonstrations via teleoperation takes 2-4 hours depending on task complexity. Fine-tuning a pretrained VLA model (like OpenVLA-7B) on this data takes 4-12 hours on a single high-end GPU (A100 or H100). Total time from "robot knows nothing about this task" to "robot performs task reliably" can be under 24 hours. Compare this to classical robotic programming: weeks to months for a skilled engineer to program equivalent behavior, with significantly less robustness and adaptability. The bottleneck is typically data collection quality (demonstrations need to be consistent and representative of the full range of environmental variation) rather than training time. What are Large Behavior Models and how do they relate to VLA models? + Large Behavior Models (LBMs) is the broader category term for large-scale models that learn behavioral policies — mapping observations to actions — rather than just processing and generating language or images. VLA (Vision-Language-Action) models are the most prominent class of LBMs. Physical Intelligence's π0 model, Google RT-2, OpenVLA, and similar systems are all VLA/LBM implementations. LBMs can be thought of as the robotics equivalent of LLMs: just as scaling LLMs led to emergent language capabilities, scaling LBMs leads to emergent behavioral capabilities — robots generalizing to novel objects, environments, and task variants that weren't in the training data. The term LBM was popularized by Physical Intelligence (pi.ai) to emphasize this parallel with language model scaling.

🤖 Physical Intelligence Lab

Four experiments: VLA input/output, imitation learning simulator, edge latency calculator, and application ROI analyzer.

VLA inference pipeline — click anywhere on the scene to place a new object

// VLA Input/Output Simulator Task instruction — Inference latency — Action confidence 7-DoF Output dimensions 30 Hz Control frequency

Task success rate vs number of demonstrations collected

// Imitation Learning Simulator Human demonstrations 50 Task complexity Medium Pretrained VLA base? Yes — Task success rate — Collection time — vs classic programming — Production ready?

Latency comparison: cloud vs edge chips — red zone = unsafe for real-time control

// Edge Latency Calculator Model size (B params) 7B Quantization INT8 Target hardware Jetson Orin — Inference latency — Memory required — Fits on device? — Real-time capable?

Cost comparison: physical AI deployment vs status quo over 5 years

// ROI Analyzer Industry Manufacturing Annual operations (tasks) 10,000 Human operator cost ($/hr) $35 — Manual cost/yr — Robot cost/yr — Annual saving — Payback period
Tags
VLA-modelsphysical-intelligenceedge-AIimitation-learningLarge-Behavior-ModelsJetson-Orinrobotic-armautonomous-roboticsprecision-agriculturerobot-learning
Share this article