VLA & GPT

51 views
Skip to first unread message

A J

unread,
Oct 5, 2026, 10:31:39 PM (3 days ago) Oct 5
to HomeBrew Robotics Club
I asked the search engine how GPT could help make a VLA model.



Strategy to study GPT‑1 → GPT‑6 for humanoid VLA:

Step 1 — Rebuild GPT‑2 from scratch
  • Understand attention math

  • Build a tiny LLM

  • This becomes your robot’s “micro‑cortex”

Step 2 — Add multi-modality (GPT‑4 style)
  • Add SigLIP/DINO vision tokens

  • Fuse them into your GPT core

  • Now you have a toy VLM

Step 3 — Add action tokens (VLA)
  • Train on robot demos

  • Predict poses, grasps, locomotion steps

  • Now you have a toy VLA

Step 4 — Add planning loops (GPT‑5 style)
  • Internal chain‑of‑thought

  • Multi‑step plans

  • Tool calls (IK, MPC)

Step 5 — Add world‑model reasoning (GPT‑6 style)
  • Predict outcomes

  • Simulate actions

  • Re-plan dynamically

This is the humanoid agentic brain.


Thomas Messerschmidt

unread,
Oct 6, 2026, 4:54:50 AM (3 days ago) Oct 6
to hbrob...@googlegroups.com, HomeBrew Robotics Club
I don’t know if it would work but I am guessing that it  would be very expensive to build. 

Of course you might make some headway via the API. 


Thomas Messerschmidt

-  

Need something prototyped, built or coded? I’ve been building prototypes for companies for 15 years. I am now incorporating generative AI into products.

Contact me directly or through LinkedIn:   




On Oct 5, 2026, at 7:31 PM, A J <aj48...@gmail.com> wrote:


--
You received this message because you are subscribed to the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this group and stop receiving emails from it, send an email to hbrobotics+...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/hbrobotics/11db2a90-063e-489b-b6ae-22aeebf3c9dbn%40googlegroups.com.

Chris Albertson

unread,
Oct 6, 2026, 3:16:49 PM (2 days ago) Oct 6
to hbrob...@googlegroups.com
If this is for your robot and you really want it to work,  I think you are starting from the wrong end.     Here is why.   I think any transformer— this means an LLM or a VLA, its job is to predict tokens.

But BEFORE you can predict tokens, you have to first make a list of all the tokens that can be predicted.     For an LLM this list is just every word in the English language.  There are only about 50,000 of them.   But for a VLA you have to invent the list yourself.

OK, you might think there is a short list of tokens all robots can do, like “grasp object” or “walk forward.”    But every robot’s “grasp” is different.  My robot has plastic fingers and yours has rubber pressure sensors, and the meaning of  “grasp” is different.  Even the meaning of “walk forward.”  What if my robot starts its walk by leaning in the direction of movement?  That would move the hands forward.   Your robot never leans first.     A VLA would need to predict hand motions differently.

I think the first step is to program or train the robot to do some primitive actions.  The first one is “stand there and do nothing,” and then “move in some direction at some speed,” and then a dozen more interesting things. Some of these primitives can be RL trained, some by teleoperation.     It really is different. If the fingers have to use the shape of the object to get under it or if there is enough friction to lift from above, motion planning would be different, and the VLA would have to have different output.

There are many details ,like exactly how the cameras are aimed, and one person writes that zebra stripe on the robot helps, and that tabletop color seems to matter too.

Without working, trained primitives, you cannot program a VLA except for some imagined generic robot.

I’m working on this too.  My solution is to assume a Unitree G1 robot.    I will never own one, but this is just an educational exercise.   Then I also want a physical robot.  So I am starting with one arm, an SO-101.  



Aside from robotics, what does it take to recreate GPT -2 from scratch?  Are you using some starting point like PyTorch or, really, from scratch and writing your own matrix math library?   Or maybe starting with a working transformer and resetting the weights to random?  It is a big job if all you do is reimplement self-attention.    Just training word embeddings seems a huge side task, but you can’t even start without that.

I just ordered some motors and electronics and an $18 camera.    Step one is connecting this to an Apple Mac mini.  Physically, it is just USB the work is the software.    


Summary:    Physical AI was where AI controls a motor in real time to perform an action.     The VLA can output tokens.  Something else has to accept tokens + the current state of the robot and environment and output a sequence of motor commands.   The critical interface is the token list.   I want to know my ideas will scale to a full humanoid, so I’ll train a simulation to stand, then walk, and I also want to know my idea can move real motors, So I’ll spend $100 on some motors.      Maybe by the end of next year I have some progress to show.    This is the way forward: robots that generalize, and we don’t hand-code 10,000 edge cases.



On Oct 5, 2026, at 7:31 PM, A J <aj48...@gmail.com> wrote:

Chris Albertson

unread,
Oct 6, 2026, 3:53:02 PM (2 days ago) Oct 6
to hbrob...@googlegroups.com


> On Oct 6, 2026, at 1:54 AM, Thomas Messerschmidt <thomas...@gmail.com> wrote:
>
> I don’t know if it would work but I am guessing that it would be very expensive to build.
>
> Of course you might make some headway via the API.
>


I’m doubting it could be done for under $10M. A group in China claimed to have beaten this cost threshold, but no one believed them.

Take just one little side task. You need a matrix that transforms embedding to Query, Key, and Value vectors. But we are talking about 1,000-dimensional vectors and a 1K by 1K matrix. That alone is 1M parameters. How long to find those? Do you search through a million steps of gradient descent or 100 million? You need to compute maybe a dozen of these. I can’t estimate the number of hours of GPU time needed. Maybe someone knows? May it is only $100 worth of electricity or maybe $100,000.

My take on this was different: What if you have the budget? You still have not connected pixels to motors. You need to walk through the story of how a photon goes through a lens and then ends up as a command sent over a serial bus.

andy

unread,
Oct 6, 2026, 8:37:13 PM (2 days ago) Oct 6
to hbrob...@googlegroups.com
I was trying to get the search engine to estimate how much energy a big robot would spend in its foundation model and guessed about $100k - 150k.

But interestingly, it said the computation and learning were the easier part, but the video data is the really big part.

It estimated that some RL-based AI would need millions of hours of video and hundreds of thousands of hours of physics simulation.

A foundation model for the Robot on the big machine would have hundreds of billions of parameters. 

I think Tesla only has 500,000 hours of video. But Nvidia can create many variations of scenarios quickly. 

The search engine seems to think the movement and validation of data was work-intensive.

Anyway, in this little thought experiment, the training is very fast, perhaps a few days.

But different levels of safety and functional regression on the AI machine would take longer.

And then, depending on the lab, days to weeks of robot testing.

But the ability to convert a 500B model to 5B on a robot frame could take months of tuning and training.

[If only we can make the hard parts easier, my guess is that robot companies use smaller models on the big machine]

Also, search indicated that the big companies would not train all the weights of the foundation model every time.

--
You received this message because you are subscribed to a topic in the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this topic, visit https://groups.google.com/d/topic/hbrobotics/eaQrucuM9Rk/unsubscribe.
To unsubscribe from this group and all its topics, send an email to hbrobotics+...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/hbrobotics/AF3D167D-9DDE-48AB-8842-C04F4F656D6E%40gmail.com.

andy

unread,
Oct 6, 2026, 8:50:19 PM (2 days ago) Oct 6
to hbrob...@googlegroups.com
Thomas,

I did some searching, and using a quantized version of an Open VLA it might take 10 - 15 hours to train.

Training on a pair of H100 might take about 1/3 the time. But with a tethered arm with vision, the 30-series GPU

would have low single-digit hertz control. So overall, it is not practical with the current hardware. But still I am 

learning, and the system is good for coding and small-scale simulations.

Best!


You received this message because you are subscribed to a topic in the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this topic, visit https://groups.google.com/d/topic/hbrobotics/eaQrucuM9Rk/unsubscribe.
To unsubscribe from this group and all its topics, send an email to hbrobotics+...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/hbrobotics/C1EBDC6E-D678-4536-95CC-16EB4A873501%40gmail.com.

Chris Albertson

unread,
Oct 6, 2026, 9:20:29 PM (2 days ago) Oct 6
to hbrob...@googlegroups.com

Yes, you would not have to “start over” and train from scratch unless you just wanted to take some new approach.    But that might be common in the beginning.




I think your AI advisor should be giving you foot notes.     There is some truth and some not so accurate stuff.

OK, to be fair, there must be a dozen ways to do each task.

Figure AI has a good web page where they explain their robot AI and give some hints

Scroll dow a little to see a disgram where they have two systems xalled system1 and system 2.    Tesla does this but runs both using a multi-headed network.bot logically similar.   Figure started a company “index” to collect data on house keeping tasks


I see videos and blogs of people training their SO-101 hand for pick and place in hours or days.    I am pretty sure it is not hard but teleoperation is tedius and slow.   But this is a not-complex primitive motor skill.



About training times, I tried looking and it is all over the map

Here is one tutorial.  And the author is more than well known, (Google "Karpathy”.) 

https://github.com/karpathy/nanoGPT.  This one trains GPT-2 (124M) from scratch on OpenWebText dataset, running on 8XA100 40GB in about 4 days. That's 8*40 = 320GB vram.

According to Meta on Huggingface, Llama-3-8b took 1.3M GPU hours with hardware of type H100-80GB. Llama-3-70B took 6.4M GPU hours.

Four days is reasonable but millions of hours?  Maybe not.










You received this message because you are subscribed to the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this group and stop receiving emails from it, send an email to hbrobotics+...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/hbrobotics/CAJZ%2Bfr0TQYRZiDTLyxfMJQHmsw12FLui-jdz6cQa_KguFuSuFA%40mail.gmail.com.

Thomas Messerschmidt

unread,
Oct 6, 2026, 9:24:19 PM (2 days ago) Oct 6
to hbrob...@googlegroups.com
Maybe I missed something, what data are you using to train it.



Thomas Messerschmidt

-  

Need something prototyped, built or coded? I’ve been building prototypes for companies for 15 years. I am now incorporating generative AI into products.

Contact me directly or through LinkedIn:   




On Oct 6, 2026, at 5:50 PM, andy <aj48...@gmail.com> wrote:



Chris Albertson

unread,
Oct 6, 2026, 9:35:02 PM (2 days ago) Oct 6
to hbrob...@googlegroups.com


On Oct 6, 2026, at 5:49 PM, andy <aj48...@gmail.com> wrote:

Thomas,

I did some searching, and using a quantized version of an Open VLA it might take 10 - 15 hours to train.

Training on a pair of H100 might take about 1/3 the time. But with a tethered arm with vision, 

Big question is train to do what?   After training what goes in and what goes out and what tasks can it do.    I’m seeing lots of different numbers but no one seems to answer those basic questions.

I looked at Open VLA it appear the model only outputs lines of text.    

andy

unread,
Oct 7, 2026, 12:06:02 AM (yesterday) Oct 7
to hbrob...@googlegroups.com
Yes, I think at this point the OpenVLA is more than my system can handle. I will continue to research 
the different ways to train robots. A key point seems to be the quality data needed to feed the pipeline.

[from search]
To create high-quality ("gold") video and demonstration data for training a robot arm on Vision-Language-Action (VLA) models like OpenVLA or SmolVLA, you need to focus on dataset diversity, synchronized multi-modal streams, and rich language/spatial annotations rather than raw video volume. [1, 2, 3]

1. Core Requirements for Gold VLA Data
Unlike standard computer vision videos, VLA models require tightly synchronized tuples at every single timestep: [1, 2, 3]
  • Visual Observations: Multi-view RGB images (e.g., third-person/overhead camera + wrist-mounted camera).
  • Language Instructions: Diverse, natural-language commands detailing the goal, paraphrases, and multi-step variations.
  • Action Trajectories: Continuous or discrete action tokens (position, orientation, and gripper state).
  • Proprioceptive State: The robot arm’s internal joint angles, velocities, and end-effector forces. [1, 2, 3, 4, 5]


--
You received this message because you are subscribed to a topic in the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this topic, visit https://groups.google.com/d/topic/hbrobotics/eaQrucuM9Rk/unsubscribe.
To unsubscribe from this group and all its topics, send an email to hbrobotics+...@googlegroups.com.

Chris Albertson

unread,
Oct 7, 2026, 12:06:51 AM (yesterday) Oct 7
to hbrob...@googlegroups.com


On Oct 6, 2026, at 6:24 PM, Thomas Messerschmidt <thomas...@gmail.com> wrote:

Maybe I missed something, what data are you using to train it.


For the stuff I care most about.  (You can’t download it because it is specific to each robot.)

  1. RL training, so there is no data, but I have to write a reward function that evaluates the behavior
  2. Imitation training, so I have to use a kind of remote control to drive the robot like a puppet and do each action a hundred of more times

For the LLM (I have no plans to use this; I’ll get a pre-trained one.)

1 There is a dataset on Hugging Face that is very widely used.  It contains 8 million documents taken from the web, put into an easy-to-use format. 37GB total.  This can be processed in about 4 days of you have 8 good GPU cards

Chris Albertson

unread,
Oct 7, 2026, 1:06:24 AM (yesterday) Oct 7
to hbrob...@googlegroups.com


On Oct 6, 2026, at 9:05 PM, andy <aj48...@gmail.com> wrote:

Yes, I think at this point the OpenVLA is more than my system can handle. I will continue to research 
the different ways to train robots. A key point seems to be the quality data needed to feed the pipeline.

That was my conclusion too.

What I am looking at now for humanoid walking is a “MLP with history buffer”. MLP is the old multilayer perceptron with an input and output layer and some hidden layers, all fully connected.    

In the top go the current robot state, joint angles, velocity, camera data, and actions come out the bottom.    But the history buffer adds most of the previous N states.  Then the model knows past joint angles back to some milliseconds in time.  Every time you run it, you move the current state to the past, and the oldest one falls off the end.     This can run faster, and like a VLA it predicts the next action given the last N states.   This is basically like time series forecasting.

Years ago, people used MLP but have moved to transformers now that they have big data centers.   But I think for a task as simple as walking, MLP should work.

andy

unread,
Oct 7, 2026, 2:11:59 AM (yesterday) Oct 7
to hbrob...@googlegroups.com
My first impression was that this was a fusion of Vision, Language, and Action (VLA).

But after searching some, it is described as a high-level policy generator.

I asked Search to explain how VLA works when the bot is asked to pick up the apple.


--
You received this message because you are subscribed to a topic in the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this topic, visit https://groups.google.com/d/topic/hbrobotics/eaQrucuM9Rk/unsubscribe.
To unsubscribe from this group and all its topics, send an email to hbrobotics+...@googlegroups.com.
flowchart_vla_Bot.pdf
bot_pickup_apple.pdf

Chris Albertson

unread,
Oct 7, 2026, 1:24:37 PM (yesterday) Oct 7
to hbrob...@googlegroups.com

On Oct 6, 2026, at 11:11 PM, andy <aj48...@gmail.com> wrote:

My first impression was that this was a fusion of Vision, Language, and Action (VLA).

But after searching some, it is described as a high-level policy generator.

I asked Search to explain how VLA works when the bot is asked to pick up the apple.


I see a terminology issue “Policy” is not what is generated.   The VLA will generate “Action chunks” which are short sequences of actions

Policy is what is trained into the VLA.  Policy is a set of rules like “to move hand forward, we move thee motors”.   I think a better way to say it is “policy generates action chunks from the current stat of the robot, the world and a command.

I think the diagram does a good job of showing WHAT happens but not. HOW it is done. The VLA wil not be built from a buch of boxes that look like the diagram.   


Before there were VLA, they used a MUCH simpler MLP to predict just one action (no action chunks)




On Tue, Oct 6, 2026 at 10:06 PM Chris Albertson <alberts...@gmail.com> wrote:


On Oct 6, 2026, at 9:05 PM, andy <aj48...@gmail.com> wrote:

Yes, I think at this point the OpenVLA is more than my system can handle. I will continue to research 
the different ways to train robots. A key point seems to be the quality data needed to feed the pipeline.

That was my conclusion too.

What I am looking at now for humanoid walking is a “MLP with history buffer”. MLP is the old multilayer perceptron with an input and output layer and some hidden layers, all fully connected.    

In the top go the current robot state, joint angles, velocity, camera data, and actions come out the bottom.    But the history buffer adds most of the previous N states.  Then the model knows past joint angles back to some milliseconds in time.  Every time you run it, you move the current state to the past, and the oldest one falls off the end.     This can run faster, and like a VLA it predicts the next action given the last N states.   This is basically like time series forecasting.

Years ago, people used MLP but have moved to transformers now that they have big data centers.   But I think for a task as simple as walking, MLP should work.


--
You received this message because you are subscribed to a topic in the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this topic, visit https://groups.google.com/d/topic/hbrobotics/eaQrucuM9Rk/unsubscribe.
To unsubscribe from this group and all its topics, send an email to hbrobotics+...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/hbrobotics/890B3CA1-1FEE-4AED-A43C-8A144FA921E0%40gmail.com.

--
You received this message because you are subscribed to the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this group and stop receiving emails from it, send an email to hbrobotics+...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/hbrobotics/CAJZ%2Bfr2aQG1w%3DkdB%3DgbYCFuCgqMFAsjwcuSZCewG%3DOoLN-fLQg%40mail.gmail.com.
<flowchart_vla_Bot.pdf><bot_pickup_apple.pdf>

andy

unread,
Oct 7, 2026, 3:35:35 PM (yesterday) Oct 7
to hbrob...@googlegroups.com
Hey Chris,

From what I understand, the VLA takes language, visual, and bot state information and generates the action command.

The policy is definitely important in RL, which is core to much of modern code. Perhaps they mean the controller block gets a high-level action like 'walk forward '. 


But I have just started to follow an RL class at Stanford,  https://web.stanford.edu/class/cs234/index.html.

The videos are on YouTube for the 2024 semester, and there are lecture PDFs. The recommended book is 'Reinforcement Learning: An Introduction ' by Sutton and Barto.  

Best!

andy

unread,
Oct 7, 2026, 4:42:53 PM (yesterday) Oct 7
to hbrob...@googlegroups.com
Chris,

Pardon me, I had not dug deep enough into the topic. It seems that MLP is everywhere.

Nearly all mainstream, end-to-end Vision-Language-Action (VLA) configurations use both an MLP and a VLA architecture. Because a VLA is a full-system pipeline rather than a single neural layer, it relies on MLPs as the "glue" that binds its distinct modalities together. [1, 2, 3]
The specific structural configurations where MLPs and VLAs explicitly operate together include:
1. The Multi-Modal Bridge Configuration (The Projector Layer)
This is the most universal configuration. Almost every single-system VLA utilizes a multi-layer perceptron right after its vision encoder. [1, 2]
  • The Setup: The visual features from models like SigLIP or DINOv2 are high-dimensional vectors. A 2-layer MLP projector sits between the vision encoder and the primary LLM backbone. [1, 2, 3]
  • How they work together: The VLA cannot process visual data natively until the MLP reshapes and mathematically shifts those image embeddings into the exact token space required by the text backbone. [1, 2]
  • Key Example: OpenVLA. It passes camera feed embeddings through a 2-layer MLP to map them straight into its Llama 2 (7B) language backbone before generating robotic action commands. [1, 2, 3]
2. Dual-System Execution Configurations
In robotic frameworks where safety and processing speed are critical, engineers decouple high-level planning from low-level muscle control. [1]
  • The Setup: This features a "System-2" slow/large VLA model working in tandem with a "System-1" fast/small execution model.
  • How they work together: The massive VLA ingests user commands and environmental images to determine what to do (e.g., "reach for the blue cup"). It outputs a conceptual trajectory token, which is then fed straight into a highly optimized, continuous MLP Regression Head that acts at a millisecond level to map those tokens directly to motor joint velocities. [1, 2, 3]
3. State-Augmented VLA Configurations
Many VLAs only read camera feeds and text, but robots also need to know where their own limbs are in 3D space (proprioception). [1]
  • The Setup: A multi-input token pipeline.
  • How they work together: Continuous data from the robot's physical sensors (like joint angles or gripper pressure forces) cannot be parsed by standard vision or text tokens. The configuration sends this raw numerical data through an MLP Encoder, which vectorizes it so the VLA backbone can process the robot's physical self-awareness alongside the camera feeds.
  • Key Example: NVIDIA Isaac GR00T-N1.6 pipelines. [1]
4. Quantized / Edge Deployment Configurations
Because running 7B to 55B parameter VLAs on a physical robot consumes too much power and causes latency, specialized hardware configurations are required. [1]
  • The Setup: Post-training weight compression (Quantization). [1]
  • How they work together: Modern configurations designed to make VLAs fast enough for real-time edge processing explicitly compress both the heavy internal Transformer attention blocks and the dense MLP layers. [1]
  • Key Example: CHASE-VLA, which applies aggressive W4A4 quantization to the MLP projections and attention components within the model to drastically lower memory traffic and power usage while preserving robotic task performance. [1]
-----------------

Reply all
Reply to author
Forward
0 new messages