Elon Musks thoughts on the Robotaxi

100% agree on this. My use of Uber/Lyft for various reason locally, on business trips and around the world on vacations has not impacted my need/want to own cars at all.
Cars last a long time, so any big change in ownership rates will take a long time to play out. A few percent of people getting rid of cars in favor of robotaxis just adds some low cost cars into the mix for lower income people to have available for car ownership.
By comparison it is NOTHING like someone deciding the cut the cord (i.e. choosing streaming instead of cable TV) because each cord cutter actually encourages more streaming content making it more attractive for others to follow them.

Mike

1 Like

About Dojo. Tesla has been very open about their AI projects looking to hire the best technical personnel available. Catch up with Tesla’s Dojo. BTW, it’s hard to BS engineers!

AI Day 2021 Aug 20, 2021
1:45:40

AI Day 2022 Oct 1, 2022
01:56:50

The Captain

1 Like

I partially agree with your point here (gobs of data) as it applies to autonomous driving. This is because it is a very safety critical activity. You need very large datasets, not because you need thousands/millions of examples of how to drive within the lines. You need all those miles of examples to have a good chance to collect dozens/hundreds (or more) examples of all the various hard-to-find edge cases, such as a construction zone narrowing to a single lane, an object falling off the back of a truck or a tumbleweed (or a rock?) rolling across the road. (You still need to augment the real edge case videos to change the lighting, weather, occlusions, more/less pedestrian traffic, etc. to give you more examples that you have no chance to actually collect live)
In the case of a robot performing an assembly task in a factory, probably not as safety critical if it drops a component or crushes a component (squeezes not hard enough or too hard) and produces a failed part every so often. If a few percent of parts have to be reworked it may still be economically viable for a robot to do most of the work. And, of course, over time, you optimize to increase yields.
But you don’t have to augment all your training data for non-optimum lighting, bad weather, pedestrians that block your path (but you still have to allow for employees accidentally getting in your way and safely avoiding them).

Mike

1 Like

This is reality vs.the instant perfection in everything. Just like not all humans need to know everything in the universe, humanoid robots also don’t need to know everything.

Think of Dojo’s size and capacity as you would think of storm drains, they have to be built to handle the largest loads but not every storm is “enormous, gobsmacking” huge. Not every AI application is either.

The Captain

1 Like

Mike, I certainly agree that a factory environment presents differences from self-driving. But I don’t think it affects the amount of data required all that much.

I’m just a layperson on the subject, but my understanding is that it’s just enormously difficult to enable computers to form a model of their surrounding environment using visual data. It’s just fiendishly complicated for them to “see” images in a useful way. That’s why visual Captcha still exists, and why even AI programs that have been trained on the entire image catalog of the whole internet still exhibit uncanny valley weirdness when generating images from scratch.

When robots are in a fixed location working with materials that are always in the same place, with identifiers that allow them to “know” exactly where everything is, they can perform amazing and complex feats of engineering and manufacturing beyond the skill of humans. And they have since the 70’s and 80’s. But you can’t build a robot that can go down to the mail room and bring a package and put it on your desk - something a human child could probably do. The tasks require different skillsets, and the latter is infinitely more complicated. Which is why all the existing companies in the robotics field - which also employ very smart engineers and AI specialists and machine learning boffins - have struggled to solve those types of problems.

And while it’s true that a robot making an error in a factory won’t result in an immediate potentially fatal collision, that doesn’t mean there’s a huge tolerance for error. It depends on the job they’re doing. Musk recently complained about the panel gaps on the Cybertruck and called for precision down to a micron, for example. If a robot incorrectly installs a wire harness at too high a rate, that’s going to affect line productivity. And obviously, when building machines like automobiles a manufacturing error can sometimes be as fatal as a driving error, depending on the system in which it occurs.

All of the above is why we’ve been able to use specialty robots for a dizzying array of tasks over the last 50 years - but a general purpose robot has been well beyond our technological capacity.

1 Like

A lot of the data (most of it??) collected by driving cars all around is redundant, is what I’m saying. In cars you are constantly collecting so that you can get a lot of the edge cases. Sometimes you know it is an edge case right away, such as when a driver has to take over. Other times (in the Tesla case or even other brands) the car is looking for certain criteria to flag it as an edge case.
It the case of training a humanoid for a factory you would not be collecting huge amounts of redundant data. Sure some would be.

As far as how much data is required to train a car vs a factory robot, for sure less is required in the factory case since it is not going to have to perform its given tasks in highly different weather conditions (rain, snow, fog) lighting conditions (daytime, nighttime, sunset glare, headlight glare, cloudy, sunny, then multiply by all the weather choices) and all sorts of other difficult situations, such as a child that is present one moment and then hidden by a hedge, etc.

Mike

3 Likes

Perhaps - but not as much less as you might think.

Again, it’s really hard for computers to “understand” visual input and translate it into useful information. And it’s hard for us to intuitively grasp that, since computers can do also do things beyond the abilities of humans. Which is the joke behind the comic at the bottom of this post, from 2014 (xkcd - can’t recommend enough). Even back then, your phone could be programmed to perform complex multivariate equations or store thousands of books - but it couldn’t tell whether a photo was of a bird or not (which even a small child can do).

So, getting back to the robots. You don’t need massive amounts of data just to make sure that the edge cases get caught. You need massive amounts of data to “teach” the robot how to construct a model of their environment using visual images. The way we “teach” a computer to recognize that an object sitting on a shelf is a cardboard box using machine learning is by feeding it massive amounts of visual data, some of which is tagged with “cardboard box” and some of which is not. Because the machine has access to that massive amount of data, it develops an algorithm for assigning which patterns of light mean “cardboard box.”

While the machine gets better at figuring out cardboard box-ness if it encounters edge cases where it’s tough to tell (like boxes wrapped in newspaper or shaped like a rhombus), it still needs the massive amounts of data just to figure out what a cardboard box looks like in the first place.

Image-based AI (like DALL-E and others) exists because there are literally billions and billions of labeled image files that people have put on the internet. So now those AI’s actually do kind of know what a bird looks like. Because they’ve seen close to a billion photos that are labeled “bird” and something other than “bird” - and they can then form an algorithm to identify what the “bird” photos have that the “other than bird” photos don’t. Nothing to do with edge cases.

So, you can’t just show have the robot (or computer) watch someone install a wire harness for a week and have it “know” what a wire harness is, or how to install it, using the processes that underlie AI. The computer won’t know what’s going on - it’s not programmed to. It has no way of knowing whether any part of the image is a car frame or a human or the harness. Only if it sees enough repetitions will it be able to develop an algorithm that distinguishes between the various objects in the scene - what parts of it are “worker,” what parts are “harness,” what parts of if are “car,” and the like. And it has to be taught what to ignore as well - if you show the computer some video of a man installing a harness in a factory with blue walls, and a second one of a factory with green walls, the computer has no idea whether that’s relevant data or not. Only if you show the computer a sufficiently large number of wire harness videos with a sufficiently varied array of backdrops will the computer be able to learn that only some parts of the image matter, while others don’t.

And that’s just for one job. If you were trying just to make a “wire harness robot,” you’d need a massive amount of data to “teach” the robot what a wire harness is and how to install it. But Musk wants to make a general purpose humanoid robot. You can’t just show it one job in one factory - you have to show it countless jobs in countless factories for it to be a “general purpose” robot, that can be put in a new factory and told to go perform some task.

You need enormous amounts of examples to train a computer, rather than program it.

1 Like

This isn’t about robotaxi (which I believe will be a huge success for TSLA) , but is it going to be really good for TSLA that they don’t use union workers and will it protect the share price? Just something to throw out there to think about…doc

1 Like

LOL so you and Albaby set up unrealistic time lines, then argue they are unrealistic, and now pat yourself on your back that you won the argument that you set up. Well done. I have to say you two are masters at setting up straw man arguments and winning every time. LOL

Andy

2 Likes

Instead of focusing on the job for a moment focus on the environment. How many orders of magnitude is the road more complex than the factory floor, the office space or the home. The humanoid robots need to be aware of the environment to get the job done safely.

An earlier post denigrated improving the design of wiring harnesses to make life easier for robots. Humans do that for humans all the time, our houses are better that caves used to be. Is that because human intelligence is lacking? It’s the pragmatic way to do things.

The Captain

3 Likes

Interesting question. In the broadest sense it’s about manipulating supply and demand which ultimately controls prices and costs. I believe the more freedom a company has to hire the more efficient it can become, all else being equal. Because of the good pay but more because of the interesting challenges, Tesla attracts the cream of the crop in such volumes that they can skim the cream of the cream of the crop.

Unions have their place to equalize the power divide between capital and labor but the more complex the job positions are the less labor needs unions, laborers can depend on their capacity to add value. Laborers with MBAs do OK without unions although fraternity houses do function like unions. :imp:

The Captain
out on a limb… :slight_smile:

1 Like

I didn’t set up the timeline for Tesla to produce the truck, Musk did. He missed it.

I didn’t set up the timeline for Tesla to produce the self driving car. Musk did. He missed it. We’re still waiting, even as other companies have self-driving robotaxis operating in some cities.

I didn’t set up the timeline for Starlink to become profitable or achieve revenue goals. Musk did, and he is at just 10% of what he promised.

I didn’t set up the timeline for hyperloop to be operating by now, Musk did. It isn’t.

I didn’t set up the timeline for the “general purpose humanoid robot”, I merely say it won’t be done in time to be meaningful to today’s investor, although someday it will probably happen.

IMO you are completely off topic here. DALL-E is a generative AI for creating images not for recognizing images or objects within an image or recognizing the actions within images temporally.

The more appropriate models for computer vision are models like Resnet50. This was the best model in the world on the Imagenet contest back around 2015 and it and its derivatives are still widely compare in informal benchmarks as well as recognized industry benchmarks for computer vision such as MLPerf. It"only" takes about 1000 images, perhaps 2000, of each of the 1000 classes in Imagenet to converge to a peak score on this model.
Imagenet has way more image classes than needed for autonomous driving having lots of different categories, such as about half different animal species including over 100 dog breeds. The dataset was designed like this to encourage the research to be very wide as well to have the ability to to discern subtle differences in details.
More appropriate for cars and robots would be an object detection model, such as SSD or Yolo. At a Tesla AI day they said they were using a model similar to Yolo (now up to version 8). This model detects 80 classes…no need to detect the difference between each breed of dog or between a dog and a wolf, for example. It also takes only 1000 - 2000 examples of each class to converge to its peak score. This object detection is what gives you the often seen examples where each object is shown with its bounding box.

The reason you need far far more data than this is that after the object detection step (on each video frame) the system has to construct and update a 3D model from it and temporally map and predict what will happen next.
Additionally, the environment that all the objects have to be reliably detected in can vary greatly as I’ve previously said (weather, lighting, occlusions, etc.)

Mike

1 Like

Thanks for the information, mschmit. I didn’t know that pattern recognition was advanced enough to require so little input for image recognition.

However, that’s not the task the robot has to solve here. Again, they’re not looking at a still photo and saying “this photo has a hammer on a table in it.” They also have to construct a 3D model from it (where, physically, is the hammer and table), so that like the car they can traverse the space. A bounding box isn’t enough - the machine has to know the exact edges of the object. The environment doesn’t change much (constant lighting), but a space filled with humans that can move anywhere and everywhere is more chaotic and unpredictable.

But the main point remains that the robot has to perform a fundamentally different task than a self-driving car. The car almost never interacts with or manipulates any objects in its environment. It’s only a slight exaggeration to say that the primary directive for a self-driving car is never to touch or be touched by any other object in the environment (save the road surface, of course). All the car has to do is move through the environment without encountering any objects. As long as it knows where the objects are, it doesn’t need much more data about them.

The robot, however, has to manipulate objects in order to be useful. It doesn’t just avoid everything. It has to grasp, turn, lift, manipulate, twist, bend, hold, crush, tear, push, or pull things - or any of a hundred other functions. To do that, it has to know whether objects are light or heavy, rigid or flexible, smooth or textured, strong or fragile, etc. It has to learn how to manipulate objects with lots of characteristics that can’t be determined visually. It needs data other than visual data to serve as training data.

Making the robot a bipedal humanoid ratchets up the degree of difficulty - because in addition to having to learn about its environment and the objects in it, it has to learn how to adjust its own form to react to those inputs. A car is enormously stable - a car at rest isn’t going to topple over. Similarly, a quadrupedal or wheeled robot has a fairly stable base as well. But if you ask a bipedal robot to walk across the room and pick up a package of unknown weight of unknown distribution, it has to do some pretty complicated adjustments to walk back with it - using data that’s completely not visual, so it can’t learn how to do it by watching lots of humans doing it.

Clearly you could make a robot that you could put in a room, look around, and say “Yes, that’s a television sitting on that table” without a lot of training data. I was clearly wrong about that. But…could you teach that robot to walk across the room, pick up the television without breaking it, and with a different balance of weight carry it back across the room, without using a ton of video training data? If it could be done with video data at all?

He didn’t miss it he just was flexible and changed the timeline. Everyone does that. Now it is here.

He didn’t miss it it is still being developed. Stay tuned.

Who cares if it is profitable. It is an amazing achievement and will eventually be profitable.

Yes it is, The Las Vegas convention center has one that operates every day.

If today’s investors you mean anyone over 70, then the chances are you are right. If you mean anyone over 20, I would say you are probably incorrect. Buffet said that the United States best times were behind it, But when you are 93 isn’t the best times always behind you?

Andy

3 Likes