Getting Claude to act more like a VLA(M)

62 views
Skip to first unread message

Thomas Messerschmidt

unread,
Oct 1, 2026, 1:42:30 PM (8 days ago) Oct 1
to hbrob...@googlegroups.com
I have been experimenting with getting Claude to act more like a VLAM. It is reluctant. It keeps telling me why it can't do what I was asking. 

I finally sent this prompt: 

Your job is to solve the problem, not tell me the difficulties.

And it began to help me solve the problem. 

Thomas   

Chris Albertson

unread,
Oct 1, 2026, 1:52:32 PM (8 days ago) Oct 1
to hbrob...@googlegroups.com
I am curious.  Can you tell us which prompt caused Claude to be reluctant?

I have only seen this when I ask for something impossible as a test to see what would happen; then it suggests that maybe what I wanted was X, Y, or Z.

Michael Wimble

unread,
Oct 1, 2026, 1:53:08 PM (8 days ago) Oct 1
to hbrob...@googlegroups.com
I had the problem once where the ai told me that other people spent years writing some software and it wasn’t appropriate for me to ask it to write the same package. I told ai much the same. What I learned at my years at Apple is that we are the experts and we do the hard stuff so people can just get on with their lives. I expect ai to do the hard stuff, not whine about it. 

On Oct 1, 2026, at 10:42 AM, Thomas Messerschmidt <thomas...@gmail.com> wrote:



Michael Wimble

unread,
Oct 1, 2026, 1:56:14 PM (8 days ago) Oct 1
to hbrob...@googlegroups.com
Not to answer for Thomas, but my request was to write a new equivalent to the groot ui program for behavior tree visualization (as opposed to several other packages called groot). It’s not even that hard, actually. The ai just assumed right away that it must be hard. 

On Oct 1, 2026, at 10:52 AM, Chris Albertson <alberts...@gmail.com> wrote:

I am curious.  Can you tell us which prompt caused Claude to be reluctant?
--
You received this message because you are subscribed to the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this group and stop receiving emails from it, send an email to hbrobotics+...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/hbrobotics/2B836EFA-6F7C-4259-8AC3-01A0FE9EC655%40gmail.com.

Thomas Messerschmidt

unread,
Oct 1, 2026, 2:06:01 PM (8 days ago) Oct 1
to hbrob...@googlegroups.com
Exactly 😁. 


Thomas Messerschmidt

-  

Need something prototyped, built or coded? I’ve been building prototypes for companies for 15 years. I am now incorporating generative AI into products.

Contact me directly or through LinkedIn:   




On Oct 1, 2026, at 10:53 AM, Michael Wimble <mwi...@gmail.com> wrote:


--
You received this message because you are subscribed to the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this group and stop receiving emails from it, send an email to hbrobotics+...@googlegroups.com.

Thomas Messerschmidt

unread,
Oct 1, 2026, 2:10:33 PM (8 days ago) Oct 1
to hbrob...@googlegroups.com
I was asking Claude how to take video directly from YouTube to train an AI. 



Thomas Messerschmidt

-  

Need something prototyped, built or coded? I’ve been building prototypes for companies for 15 years. I am now incorporating generative AI into products.

Contact me directly or through LinkedIn:   




On Oct 1, 2026, at 10:52 AM, Chris Albertson <alberts...@gmail.com> wrote:

I am curious.  Can you tell us which prompt caused Claude to be reluctant?
--
You received this message because you are subscribed to the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this group and stop receiving emails from it, send an email to hbrobotics+...@googlegroups.com.

Chris Albertson

unread,
Oct 1, 2026, 3:45:25 PM (8 days ago) Oct 1
to hbrob...@googlegroups.com


> On Oct 1, 2026, at 10:52 AM, Michael Wimble <mwi...@gmail.com> wrote:
>
> I had the problem once where the ai told me that other people spent years writing some software and it wasn’t appropriate for me to ask it to write the same package.

But we all know that the LLM was trained on available text. It said what an expert wouild say.

In other words the AI is so good at predicting what an intelegent person would say that we begine to think it is an intelegent person. But really it is only copying.

Still the better response would have been “writing that package is beyond my current abilty.”

Chris Albertson

unread,
Oct 2, 2026, 1:27:02 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com
I think this is a reasonable request.   You were not asking Claude to use YouTube to train an AI; you were asking how it could be done.   I just tried this using Google’s Gemini.

Using this exact prompt:  Please make a plan I can follow, I want to take video directly from YouTube to train an AI.

The AI wrote a very nice and actionable plan.  I read it, and there are only one or two mistakes that don’t make sense, but they are minor.    But here is the difference:   I asked it for a PLAN, not to actually do it for me.    When you see the plan, you will see that it would be impossible for any current AI to do this.   Then, after writing all the steps, I should follow; it asked for more information and offered to help more.     Perhaps we should all be using Google’s Gemini when what we want is a teacher.

I find that the best words to place in front of a prompt are “Please explain how I can do X” or “Create a plan for doing X”.” And then, later, I might ask for help if one of the steps is not obvious.   When you tell the AI to solve the problem, it can’t.  Ask it for a plan or for advice on how the problem might be solved.

Thomas, I remember you posted a photo of a Robot arm some days ago.  Frankly, the photography was not good, and I was not impressed.   But later I found it was an SO-101 arm.  (the AO-101 is open source) Now I am impressed, and I want one.  They are cheap if you have a 3d printer, under $100 each, and I’d go for the 12-volt servos.    I am downloading the CAD (.step) files today.    Now that I know what an SO-101 is, I am revisiting your email.

My plan for the SO-101 is to first use RL training in simulation.    This can run with no human input to train the grasp-and-lift task.    Human input will be required for tasks like “fold a T-shirt.”

Read the plan below.  Using YouTube is a LOT OF WORK.  You have to edit the video and just take the good stuff and label it carefully. With phases like “using a knife", "tying shoelaces",  or whatever.  Many YouTube videos have long talking head segments, and you have to cut that part out.   You can spend hours with a video editor.      I would not start with YouTube, us a video camera and a very simple task, like “pick and place"

The plan also (big mistake, I think) left out “data augmentation,” where you use different crops and rotations on the same video to make 4 or 8 clips from each clip and “augment” the amount of data.  This makes the training see action, not pixels, and helps prevent overfitting.   So do not blindly trust the AI, this was a mistake, not fatal, but you have to read and verify everything the AI tells you.




Here is the plan for using YouTube video: (but you can already re-create it by copying/pasting my prompt.)

🗺️ Step-by-Step Training Plan
1. Data Collection and Extraction
To train an AI, you first need to securely download and store the raw video files or their corresponding audio/text components.
  • Extraction tool: Use open-source command-line tools like yt-dlp to safely download videos, audio tracks, or subtitles.
  • Metadata saving: Extract and save accompanying video data (titles, descriptions, tags, view counts) into a JSON or CSV format.
  • Automation script: Write a Python script using the YouTube Data API to fetch a list of URLs based on specific search queries or channels.
2. Data Preprocessing
Raw video files are far too large and unstructured for an AI model to ingest directly. You must clean and format them.
  • Video frame sampling: Use libraries like OpenCV or FFmpeg to convert videos into sequential images (frames). You usually only need 1 to 5 frames per second rather than the full 30+ FPS.
  • Dimensionality reduction: Resize all images to a uniform resolution (e.g., 224x224 or 512x512 pixels) and normalize pixel values.
  • Audio processing: If training an audio or speech AI, convert the audio track into waveforms or Mel-spectrograms using Librosa.
  • Text tokenization: For transcripts, clean the text and tokenize it using tools like Hugging Face Tokenizers.
3. Dataset Annotation and Structuring
Your AI needs to know what it is looking at or listening to.
  • Labeling: Match each processed video clip or frame with its corresponding label (e.g., automated transcriptions for speech, or category tags for computer vision).
  • Storage format: Save your structured data into machine-learning-friendly formats like TFRecords (TensorFlow) or WebDataset (PyTorch).
  • Dataset splitting: Partition your data into 80% training, 10% validation, and 10% testing sets to prevent overfitting.
4. Model Selection and Training Setup
Set up the environment and architecture where the learning happens.
  • Framework selection: Choose a deep learning framework like PyTorch or TensorFlow.
  • Architecture selection: Use a model suited for video data, such as a Video Transformer (TimeSformer), a 3D Convolutional Neural Network (3D CNN), or a multimodal model like CLIP if matching text to video frames.
  • Compute infrastructure: Ensure you have access to dedicated hardware, such as an NVIDIA GPU or a cloud-based equivalent (Google Colab, AWS EC2, or RunPod).
5. Training, Evaluation, and Fine-Tuning
Feed the data into the model and monitor its progress.
  • Compute metrics: Track loss, accuracy, or mean Average Precision (mAP) during training using tools like Weights & Biases or TensorBoard.
  • Validation testing: Regularly test the model on your validation dataset to make adjustments to hyper-parameters (like learning rate and batch size).
  • Final deployment: Export the trained weights for inference testing on your independent test set.

To help tailor this pipeline to your project, could you tell me:
  • What type of AI model are you trying to train? (e.g., object detection, video generation, speech-to-text, action recognition)
  • What programming language or framework do you prefer to use?
  • Do you already have a list of target YouTube videos, or do you need help writing a script to find them?
AI can make mistakes, so double-check responses

daniel miller

unread,
Oct 2, 2026, 4:32:03 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com
I spent several months going down this rabbit hole (claude as only NN). The simple fact is that Claude lacks intuition about our 3 dimensional world. It's like talking to a blind person about how beautiful a sunset is. They can interact with you, even use the vocabulary (Stevie Wonder: "You Are the Sunshine...) but in the end they just don't have first-person experience and so make subtle mistakes.

Not to mention, the top models are too damn slow.

Chris Albertson

unread,
Oct 2, 2026, 5:36:40 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com
Yes, of course.  Any online LLM is going to be 1000 times too slow.

If you think LLMs are not the way forward for robotics, you are in good company.  I think this is now mainstream opinion.

Just think about an arm,….   The commands to the motors need to be updated 30 times per second, once per video frame.  You would hope the AI that is controlling the arm can run end-to-end 30 times per second.

A humanoid robot ideally would have 100 Hz updates or even 200 Hz to the legs and body.

The only thing Claude could be used for is the very highest level planning where you say “Robot, clean up the kitchen and make it look neat and clean.” Then Claude writes a plan that some other AI executes.

Hierarchical control seems to work.  You have a real-time layer that does simple tasks and a slower but smarter layer over that, and possibly a third, maybe.

But as I wrote, the best use of large online LLMs is as teachers.  Ask them, “How should I approach this problem?”   


Everyone here in this thread seems to have missed the point that motors really do need to be updated about as fast as video frames.   Both for that same reason: smooth motion.   I just do not think a normal person is going to run any transformer model at 30 Hz.  Even with an expensive GPU.  You have to go for something simple if it is to be that fast.



Dave Everett

unread,
Oct 2, 2026, 6:00:18 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com
On Sat, 3 Oct 2026 at 03:26, Chris Albertson <alberts...@gmail.com> wrote:

Thomas, I remember you posted a photo of a Robot arm some days ago.  Frankly, the photography was not good, and I was not impressed.   But later I found it was an SO-101 arm.  (the AO-101 is open source) Now I am impressed, and I want one.  They are cheap if you have a 3d printer, under $100 each, and I’d go for the 12-volt servos.    I am downloading the CAD (.step) files today.    Now that I know what an SO-101 is, I am revisiting your email.

I just bought 24 of the ST-3215 servos they use, 12v 30kg-cm, for under $700. There are good deals around on Aliexpress.

The startup Vulcan Robotics I mentioned a while back is using these arms and training system.

Check out Illia's youtube channel for good information on using this with Pi 0 and Pi 0.5 VLA training tools:

Dave

Thomas Messerschmidt

unread,
Oct 2, 2026, 6:08:35 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com
I see your point. I am still working through it. Thanks for the insight. 



Thomas Messerschmidt

-  

Need something prototyped, built or coded? I’ve been building prototypes for companies for 15 years. I am now incorporating generative AI into products.

Contact me directly or through LinkedIn:   




On Oct 2, 2026, at 1:31 PM, daniel miller <danb...@gmail.com> wrote:



Thomas Messerschmidt

unread,
Oct 2, 2026, 6:27:17 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com, hbrob...@googlegroups.com
Actually it is composed of 4 AX-12 Dinamixel servo motors. Not High Wonder. 



Thomas Messerschmidt

-  

Need something prototyped, built or coded? I’ve been building prototypes for companies for 15 years. I am now incorporating generative AI into products.

Contact me directly or through LinkedIn:   




On Oct 2, 2026, at 3:00 PM, Dave Everett <daveev...@gmail.com> wrote:


--
You received this message because you are subscribed to the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this group and stop receiving emails from it, send an email to hbrobotics+...@googlegroups.com.

Chris Albertson

unread,
Oct 2, 2026, 6:53:38 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com

I think there must be some miscommunication.   Which arm is this?     I thought it was SO-101.    



Assuming it is SO-101 we can look at the github repository and see the bill of materials.    It reads “FEETECH,” and not HiWonder and not Dynamixel.


So is there another arm?   I’d actually like for find more.  The SO-101 is only 5-dof in the arm plus one more in the gripper.  The math works out so that we can not control the approach angle to the object being gripped

Thomas Messerschmidt

unread,
Oct 2, 2026, 7:10:18 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com, hbrob...@googlegroups.com
My arm is a Dynamixel. It is made from 4 servos, Dynamixel hardware , and a claw. I added the base. 

It was given to me by a friend. 


Thomas Messerschmidt

-  

Need something prototyped, built or coded? I’ve been building prototypes for companies for 15 years. I am now incorporating generative AI into products.

Contact me directly or through LinkedIn:   




On Oct 2, 2026, at 3:53 PM, Chris Albertson <alberts...@gmail.com> wrote:


--
You received this message because you are subscribed to the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this group and stop receiving emails from it, send an email to hbrobotics+...@googlegroups.com.

Thomas Messerschmidt

unread,
Oct 2, 2026, 7:28:00 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com
image0.jpegimage1.jpegimage2.jpeg



Thomas Messerschmidt

-  

Need something prototyped, built or coded? I’ve been building prototypes for companies for 15 years. I am now incorporating generative AI into products.

Contact me directly or through LinkedIn:   




On Oct 2, 2026, at 3:53 PM, Chris Albertson <alberts...@gmail.com> wrote:



daniel miller

unread,
Oct 2, 2026, 7:30:02 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com
Of course we didn't try to servo with the LLM - but we did try planning and related higher-level functions. Our thinking was to get POC even if 10X too slow, then assume we could find a smaller, faster bot (we looked at gemma3.2). 

It does make me suspicious of VLA's -- where do they draw the line?

Would a model like jev be of use here?

Chris Albertson

unread,
Oct 2, 2026, 7:33:07 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com


> On Oct 2, 2026, at 4:09 PM, Thomas Messerschmidt <thomas...@gmail.com> wrote:
>
> My arm is a Dynamixel. It is made from 4 servos, Dynamixel hardware , and a claw. I added the base.
>
> It was given to me by a friend.

After I read “Dinamixel,” I figured it as a Koch V1.1 wich is very much like the SO-101 but the Koch has Dynamixels and the SO-101 has Feetech. Both servos use a serial interface and have kind of the same features, although Dynamixel has a better reputation for quality.

But with 4 DOF, your arm will be able to control one claw angle. The SO-101 is 5-DOF and can control 2 angles. Ideally, you want 6-DOF and the ability to control the full roll, pitch, and yaw of the claw. But if the goal is only learning to train arms and not to actually fold T-shirts, then 4-DOF is fine; you simply restrict the problem domain to symmetric objects where the claw angle does not matter.

Another thing that is needed is a camera. I read that many people are using the so-called 32x32 camera as a wrist-cam. 32x32 refers to the size of the camera’s PCB mounting holes. They are USB cameras that cost about $20 each. It seems you want one on the wrist and one mounted high for a bird’s-eye view. Next, the upgrade path is to add a second arm with its own wrist camera. Then train it on two-handed jobs.

The first step is to get either arm working with a Python API. Can’t train without that.


Thomas Messerschmidt

unread,
Oct 2, 2026, 7:36:12 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com
Almost any LLM will be too slow. It would probably need to be a combination of technologies, and more than one AI system. 

Of course the use case will determine the technology. A fully functioning humanoid robot that emulates people  would be a whole lot different than either having a robot arm lift a bottle or having bipedal robot legs balance dynamically, and walk.



Thomas Messerschmidt

-  

Need something prototyped, built or coded? I’ve been building prototypes for companies for 15 years. I am now incorporating generative AI into products.

Contact me directly or through LinkedIn:   




On Oct 2, 2026, at 4:29 PM, daniel miller <danb...@gmail.com> wrote:



Chris Albertson

unread,
Oct 2, 2026, 7:59:18 PM (7 days ago) Oct 2
to hbrob...@googlegroups.com


On Oct 2, 2026, at 4:29 PM, daniel miller <danb...@gmail.com> wrote:

Of course we didn't try to servo with the LLM - but we did try planning and related higher-level functions. Our thinking was to get POC even if 10X too slow, then assume we could find a smaller, faster bot (we looked at gemma3.2). 

It does make me suspicious of VLA's -- where do they draw the line?

Would a model like jev be of use here?


I’m looking at how others have solved this.    The best description is for Figure AI.   They talk about a layered system.  But then they are not so constrained by budget, and they just got another billion dollars of funding.    For us normal people, I think the only possible way is to make the lowest-level model VERY simple so it can run at 20Hz using a computer that costs well under $500.    That is not a magic number; $500 is simply my upper limit.    $100 might be possible.

So, I would use a simple “multi-layer perceptron” as the real-time model.  This is as simple as it gets.  The input layer has the input command and the current device state and maybe the prior state(s) plus the current processed and normalized video frame.    The output layer is simply the numbers that are to be sent to the motors (whatever your motors need: torque, position, velocity, or whatever). Then some hidden layers, how many depending on what you are doing.

Then you train the above to do some small set of commands like grasp, lift, set down, move, and whatever.    You need a set of primitive commands like that.

Then your VLA is trained to predict the next primitive command, and you run the VLA once per each primitive command.  Its input would be a medium-level command and a video frame from each camera.  Medium-level commands might be “ move object from countertop to cupboard.”    And then you run the VLA until it predicts an “end” token.

But nothing works until you train the robot to do some library of primitive functions.  I assume you use RL training because the primitives are so simple you could write a loss function.   But imitation learning might work.  Or both, but I admit I don’t know how to combine imitation and RL.   The trouble with imitation is that you (the builder) need to perform the action maybe 1,000 times.   RL can run while you are at work or sleeping.

Definatly start with a simulation.    

James H Phelan

unread,
Oct 3, 2026, 10:42:04 AM (7 days ago) Oct 3
to hbrob...@googlegroups.com

Thomas, et al.

I keep seeing the SO-101 arm pop up as a popular choice and compatible with LeKiwi which some of our members have.  I asked Gemini about it:

Customer service for Hiwonder is generally described by the robotics community as a "mixed, do-it-yourself" experience. While they are a legitimate manufacturer with functional hardware, their post-purchase support has notable limitations. [1, 2, 3]

Support Channels & Response Times


  • Primary Contact: Official support is managed primarily via email at sup...@hiwonder.com or through their phone line. There is no native website live chat or immediate ticketing system. [3, 4]
  • Communication Barrier: Because Hiwonder is based in Shenzhen, China, customers in North America or Europe frequently report a one-day delay on email responses due to time zone differences. Language translation gaps can sometimes make complex troubleshooting slow. [5, 6]
  • Official Forum: Hiwonder maintains an Official Technical Support Forum where agents respond publicly to user issues like connection errors, flashing firmware, and missing code dependencies. [7]

Community Consensus


  • The Good: Customers report that physical kits ship well-packaged and rarely arrive with missing parts. If your problem is a known software bug, they usually have a tutorial or user manual to cover it. [8, 9, 10]
  • The Bad: For complex or non-standard hardware issues, users on Reddit report that "you are mostly on your own" once you buy. The documentation can sometimes feel disjointed or hard to navigate. [2, 11]
  • The Silver Lining for the SO-101: Because the SO-101 arm is deeply tied to the Hugging Face LeRobot framework, you do not have to rely solely on Hiwonder. The open-source LeRobot community on Discord and GitHub is incredibly active and often provides faster, more comprehensive technical troubleshooting than Hiwonder’s internal staff. [1, 12]

If you are dealing with a specific issue, let me know:

  • Is your problem hardware-related (e.g., servo errors, assembly)?
  • Is it software-related (e.g., Python scripts, LeRobot connection, environment preparation)? [5, 7, 8, 13]

I can help point you toward the exact community fix or documentation path! [1, 8]

James

James H Phelan
"Nihil est sine ratione cur potius sit quam non sit"
Leibniz
Reply all
Reply to author
Forward
0 new messages