Getting out of the way: my robotics crash course (thisismypersonalblog.com)
63 points by systemerror 2 days ago
jvanderbot 7 hours ago
First law of calibration: Make sure you do an error analysis.
Finding the table surface is pretty useless using a top-down view, even with April tags, because the range error to an april tag is much more than the bearing (pixel) error to the april tag. You basically have trouble observing the thing you're trying to measure.
If you do this again, ask your agent to conduct this analysis and make sure your desired calibration variables are observable with small error. A second camera from a 45deg angle or even on the table would go a lot further, but then of course other things become unobservable.
Nice workaround using a proxy for force sensing to get touch info, however. And neat project overall!
systemerror 5 hours ago
This is very useful info! I've added the additional 45 degree camera and so far it's seems like there is a regression on skill but I think it's because it wasn't available in the earlier parts of the training.
Can you elaborate on the calibration variables comment? This sounds useful but how would I apply these observations?
jvanderbot 5 hours ago
This is long and rambly - let's email (see profile) if you want to work through it together.
You're effectively trying to understand how motor inputs change the end hand position, and in particular, you want to know where the table top is so you can position the hand close to it to pick up/ put down.
This means you have some tuning to "learn" before you can apply a control policy / algorithm - and you should be careful how you phrase this so claude/ai can pick up the right vocabulary and bias towards good solutions.
Adding multiple views helps as follows:
0. Measure from multiple views the table top - Keep cameras stead and rigid, and ask claude to use opencv to do multi-view registration so the plane of the table is known precisely. Keep the cameras steady throughout this process - if they wiggle, you can do multi-view registration each measurement...
Paint the "finger tips" bright orange. Not kidding. Use a very flat chess board under the arm for your "workspace". Also not kidding.
1. Move arm to known position, the multiple cameras will measure the april tags movement AND THE FINGER TIP LOCATIONS. The chess board provides a very nice texture. Or a big flat texture of any kind helps here. Since we know the cameras and table positions, we're getting closer to knowing how the arm movements move the hand w.r.t. the table.
2. If your arm has encoders, then manually touch the table at several points, and claude will record the april tags + joint angles + finger locations.
2b. If your arm does not have encoders, then manually touch the table at several points, and record the april tags + finger locations only, but SPECIFY that the arm is now touching the table. --> Claude can use more opencv code to actually locate the touch point on the table. The measurement is noisy, but you will use many movements to figure it out.
Repeat many times, 10-20, multiple touch points. Then, ask claude to do an error analysis and suggest more touch points. Tell it to use system identification techniques / camera/arm calibration techniques. tell it to research these techniques and report errors.
When you're done, you have to do repeatability experiments - the key is this is now automatic. Claude picks joint angles for the arm, the fingers move, the april tags move, the multiple camers calibrate and record positions, and claude repeats. The manual steps are just the bootstrap - you should be automated now.
What you want is claude to output and ORDF of the whole system - bang you can now control the arm using off the shelf software, which claude is happy to set up for you.
When it comes time to pick things up, the multiple views will locate the object precisely and a control network or algorithm can plug in to control it via the ORDF.
jacquesm 3 hours ago
WoodenChair 3 hours ago
I thought this was going to be a post about learning robotics.
It seems more like a post about hooking up an LLM to a pre-made robotic arm.
While that's interesting, I wouldn't have labeled the post "my robotics crash course."
Unfortunately, I think this is in some sense another example of LLMs substituting for actual learning. While I'm sure the author is learning something about robotics from doing these experiments, I doubt it's as much as he would have gotten from say reading a couple chapters of an introductory robotics for dummies book.
systemerror 2 hours ago
I think my approach to learning most things I'm interested in is to set a basic goal and learn exactly enough info to get me to that goal. If the goal is too ambitious, I'll set a more reasonable goal and try again. If I'm able to achieve the goal, set a more ambitious goal and learn more during that process. Trying to learn everything about robotics is not a reasonable goal at this time so I'm doing my best to set myself up for success.
pessimizer an hour ago
The only problem with this approach for learning is that LLMs are meant to be something that you give goals, and it solves the problem. There's no space for learning there. The challenges become stating the goal precisely, and working around the stupidity and idiosyncrasies of the LLM.
Setting tiny goals, getting to them yourself, then setting a slightly larger goal, that's much more intense.
What I mean is that I'm not sure what "success" means in this context. There are already programs that control robot arms. I would think that success would mean that you managed to write a program that would control a robot arm.
You're an experienced programmer although maybe not a physics or mechanical person. Executing this would mean that you learned the mechanics (most people trying this don't have your programming experience, and have to try to do both things at once!)
Learning the mechanics would mean that you would have a good instinct to critique the AI when success in future projects would mean manifesting something you'd never seen before (and being happy you had AI to help you punch above your weight.) This would also give you a good foothold to understand more complex movements and coordination.
At least that's my perspective. The deliverable isn't on the table, it's in your brain. Asking the LLM to do it is like asking the LLM to copy a famous painting. You copy a famous painting so you can learn the movements of the person who painted it, not for the painting itself.
systemerror a minute ago
jvanderbot 2 hours ago
True in the sense of learning fundamentals from the bottom up.
Only partially true in the sense of learning how the system might work and doing problem-directed learning.
Eventually, all the problems you'll encounter getting an arm automated are self-discovered, and if you stick with it, you'll learn all you need eventually.
So this is basically step 0.
pj_mukh 7 hours ago
Shoutout to people who don't overthink blog setups. "thisismypersonalblog.com" AND GO. It'll do the job.
On a technical note, your setup is not far from where professional setups are headed [1]. Astra is really doing an end-run around (for now) research Robotics setups.
[1]: https://x.com/ihorbeaver/status/2104646736652447854?s=20
jvanderbot 5 hours ago
I would rate that as an AI-assisted setup for a traditional research robotic workspace. No different than "Claude set up ci/cd for this project", just much more hands-on (pun intended). Software ate the world so anything that can generate software is going to make world-space interactions easier.
bambataa 5 hours ago
I also got a SO-101 and mucked about a bit. I stopped partly because I realised my idea of putting a laser on the end of it had potential risks but also when you’re just prompting an agent to do things it’s hard to feel that you’re really “learning robotics”.
This post has given me some inspiration though!
thallavajhula an hour ago
Love your domain name!
ainch 5 hours ago
Curious to see how they get on with VLAs. They sound great till you have to sit and record hundreds of teleop demos to teach them how to solve your task...
Mystery-Machine 4 hours ago
Anyone tried using Jev for fast AI robot decisions? Would it make sense?
thomasikzelf 3 hours ago
What kind of decisions though? jev does not output joint angles, and if it did it also needs to know how. It also does not take in images.