Perhaps because I have experience in the airline industry, and know that they’ve been using multiple sensor fusion across redundant and divergent inputs since the 1980s? (Airbus was the first, Boeing started within a year.)
Sure, and there are over 20 companies trying to crack autonomous driving. Every single one of them except one “has an opinion” that doing so by vision only is a dead end, because it can only get you so far. Meanwhile Tesla has promised 10 million purchasers that their car can drive autonomously the way they bought it, so their “opinion” has billions of dollars in liability if they decided wrong a decade ago.
Legal liability for statements and promises made to purchasers.
If I invert your analysis, if Tesla gets it right, it means 10 million potential new autonomous subscribers and a moat vs the competitors.
As an investor, I assume the business is operating in good faith and in the best interest of shareholders and mindful of its customers and their safety. I understand the argument you’re making. Neither one of us knows with any certainty the motivation of the company is anything other than what it claims to be.
There is no doubt this is a high stakes gamble. I understand the rationale of anyone that puts capital elsewhere. One of my favorite Investors, Charlie Munger, had said he would never bet with or against Musk. As obvious as you may think the error is in his approach, the future is not set. Even if he has it wrong, Tesla functions so much like a start up, it may simply pivot to what will work.
Yes, the people without a scaling autonomous product do say that.
I didn’t write that vision-only is “flawed”:
I wrote that it may be less capable than multi-sensor for super-human safety.
And I wrote that this rationale is flawed:
Because, according to Waymo (and other experts):
Here’s Nuro:
I don’t know? Cost? Ignorance? Both?
Five minutes of a web search will find answers on how and why multi-sensor works in machine learning and while conflicts among sensor data is a thing that happens, it is nowhere near being some kind of dealbreaker in modern models - as Waymo’s millions of weekly autonomous miles has shown us for years now.
How many EVs can one buy for the price of a jet liner?
GoogleAI:
You can buy between 2,300 and 6,000 standard electric vehicles (EVs) for the price of a single new commercial jetliner.
The exact quantity depends entirely on whether you are looking at a smaller single-aisle aircraft (like an Airbus A320neo) or a massive long-haul widebody aircraft (like a Boeing 787 Dreamliner). [1]
How is that a moat? If LiDAR gets in the way, they can turn it off. If it doesn’t, it’s a $200-$500 addition in hardware addition to a $30,000 vehicle. Hardly seems like a death sentence to me. Might pinch a point or two from the margin, but then it might also be the thing that cracks the puzzle. (And so far, it is.)
PS: If it’s “so hard” to combine the disparate sensors, how come Waymo and Zoox are already doing it?
This one’s easy. He’s convinced it will work. And he is very much a risk-taker, willing to do things that everyone says are the wrong choice. It’s why Tesla still eschews advertising: it’s insane that a company that largely makes a consumer product doesn’t have personnel working to protect and manage their brand, but that’s an idee fixe for Musk, and he’s not going to back off it. It’s why the Cybertruck looks the way it does. Musk isn’t going to listen to people telling him that virtually every consumer-facing company does advertising, or that it’s a good idea to design a car to conform to consumer expectations and preferences for car aesthetics.
And Apollo Go. They don’t operate in the U.S., but they’re the biggest robotaxi provider outside the U.S. (mostly in China) and have already scaled to 27 cities. Their operations are roughly the same size as Waymo right now (~350K weekly rides to Waymo’s ~500K). They seem to have found a way to combine disparate sensors, since their vehicles use LIDAR and hi-def radar in addition to vision.
But, this is part of the reasoning Musk gave for simplifying the process and going with vision only.
If it’s okay with you, I’m going to reserve final judgement until we see whether Waymo can deliver a cost effective service or Tesla can do so with its current approach.
In fact it was Andrej Karpathy, the head of AI at Tesla at the time, who promoted this idea:
GoogleAI:
Andrej Karpathy argued that autonomous vehicles do not need LiDAR, promoting a camera-only, vision-based approach modeled on human biological driving. [1, 2]
Core Arguments on Cameras vs. LiDAR
Sufficiency of Vision: Karpathy stated that because human roads and traffic systems are explicitly designed for visual consumption by human eyes, cameras provide all the necessary information and high-bandwidth data required for driving. [1]
The Recognition Problem: During Tesla’s Autonomy Day, Karpathy noted that LiDAR struggles to perform critical semantic identification—such as distinguishing between a harmless plastic bag and a dangerous rubber tire on the road. LiDAR yields active point clouds but bypasses the fundamental semantic problem of visual recognition. [1, 2]
No High-Definition Maps: Systems relying on LiDAR often depend on pre-mapped, centimeter-accurate High-Definition (HD) maps. Karpathy argued that maintaining massive HD map databases creates a brittle geographical dependency that humans do not rely on, whereas a pure vision neural network can navigate novel intersections dynamically. [1, 2]
Cost and Complexity: Karpathy and the Tesla team viewed redundant active sensors like LiDAR and radar as organizational and engineering liabilities that add cost, supply chain friction, and noise to the neural network training pipeline. [1, 2]
If I have to pick between Fools and an AI expert…
GoogleAI:
Andrej Karpathy is a world-renowned Slovak-Canadian AI researcher, engineer, and educator who currently works on the pretraining team at Anthropic. He is widely recognized for his foundational contributions to deep learning, computer vision, and large language models (LLMs). [1, 2, 3, 4, 5]
Career Milestones
Anthropic: Joined the frontier LLM pretraining team.
Eureka Labs: Founded an AI-native education platform.
OpenAI: Served as a founding research scientist.
Tesla: Acted as the Senior Director of AI and Autopilot Vision.
Academic Foundations: Main architect and instructor of Stanford’s landmark deep learning course, CS231n. [1, 2, 3, 4]
Key Concepts & Innovations
Agentic Engineering: Championed a shift from “vibe coding” (loose prompting) to highly structured, multi-agent frameworks. [1, 2, 3]
AutoResearch: Designed autonomous workflows aimed at eliminating human bottlenecks in model optimization. [1]
The “Three-Layer” Framework: Developed a rigorous methodology utilizing a structured spec, an AI verifier/critic, and an optimized environment to build complex software. [1]
Complex LLM Benchmarking: Pushed the boundaries of AI capabilities by tasking frontier models with complex, autonomous code generation, such as procedurally generating a 3D Lord of the Ringsenvironment. [1]
What is so odd to me is the absolute certainty with which these Fools speak on four of many unknowns.
Vision only will never work or the lack of redundancy forever creates unacceptable risks
Waymo’s approach will at some point reach a unit cost that is in fact profitable.
Tesla and really Musk’s motivations are simply callous cost considerations or that they are simply too far down the rabbit hole to turn back.
Even if Tesla solves the issue of vision only for full autonomy, somehow, Waymo will at scale with the additional systems deliver a product and service at equal margins.
You don’t have to, any more than we’re picking between Waymo’s AI experts and you when we respond to your quoting of Andrej Karpathy. We’re pointing out what Waymo’s folks are saying - that if you only have a single sensor input, you hit a lower ceiling down the March of Nines. You don’t have to believe the Waymo experts, but they also are fairly accomplished in their field. While resolving multiple sensor inputs certainly requires more effort, and may make early progress much slower, there’s no doubt that several independent companies have managed to solve that problem: Waymo, Zoox, Apollo Go.
Most of the arguments in favor of limiting your sensors to just vision have disappeared. The cost of LIDAR has collapsed. High-def radar is now cheap and ubiquitous. The perceived technical difficulties of integrating multiple sensor inputs have not manifested - autonomous vehicles with multiple sensor arrays have autonomous safety records that are vastly better than humans.
There’s no more reason to eschew other sensors today than there is to limit yourself to just two cameras because humans only have two eyes. Yes, it is possible for driving to occur with only two vision inputs - but that doesn’t mean that’s the smart way to tackle the problem.
The question isn’t between Fools and an AI expert, it’s between your Tesla expert and every other company with dozens of other experts all saying “using multiple sensors is better.”
I don’t see anyone saying “vision only” will never work. Just that “vision only” is not working now, while others are pursuing multiple sensors and it is working . Perhaps “vision only” will work a year from now. Or five years from now. Or 50 years from now. We are just able to say “it doesn’t work now” because, uh, it isn’t working now.
This seems likely, given that both systems will require fantastical compute, but one requires the addition of perhaps an additional $200 in hardware.
Tesla’s advantage is that they already produce cars in quantity, so if their system works they can win. Waymo’s advantage is that they have a system which is functioning now, but must be grafted onto somebody else’s platform. If “vision only” doesn’t work (timely) then Waymo wins. If it does, Tesla has a great opportunity - but it’s still not the same opportunity they were talking about a couple years ago: that there would be millions of Tesla’s (bought by others) capable of instantly starting an Uber-killer, and being a toll-taker for having produced an app which connects riders with owners’ cars and have 30% or better margins.
Now the model is “We will produce the cars, we will own the cars, we will set up maintenance and service shops city by city, we will run the business, operate it like Hertz, and have taxi-cab margins.
The current Tesla has 9 cameras, never get tired, never distracted, and never is under the influence. In many ways, even Supervised FSD, may already be improving safety for everyday drivers.
So what? The same is true of Waymo’s vehicles, as well as Apollo Go and (I presume) Zoox. I mean, not the exact number of cameras, but the general point.
FSD is a great ADAS system - and like many ADAS systems, it certainly improves how human drivers perform. But that is a very different proposition than being capable of driving autonomously.
Which, again, was the point of Waymo’s presentation. If you limit your range of sensor inputs, you can make more rapid progress towards a really good ADAS. One that has the features you attribute to FSD. But if you limit yourself to a single sensor, you move more quickly to a lower ceiling of capability - and your really good ADAS can’t continue progressing into being an actual autonomous driver.
Pointing to how awesome FSD is as an ADAS system isn’t a counter to this point. Especially since Waymo and Apollo Go have now pretty firmly established that multiple sensor inputs are not a technical bar to developing a safe AI driver.
You know what else they have? Hearing. If they hear a siren they “decide” if they have to move out of the way. Coming from behind? Yes. Siren on an adjacent street? Probably not. Siren and flashlight light in the camera view? Yes.
Imagine that: combining multiple sensor inputs to make driving decisions. I’m told that’s practically impossible.
I’ve even seen noted experts say it’s too hard. A dead end. Useless. Unnecessary. Go figure.
Again, you’ve made my point. We do not know the final outcome but these were prior comment on the thread. It sure sounds like he/she is saying it will never work.
“Yep. Waymo’s position is that if you go with vision-only, you don’t have to put in the work and effort to solve the decision-making inputs for multiple inputs. That lets you go faster, and you get to a really good ADAS more quickly. But because you’ve taken the shortcut, you can’t get down the March of Nines to an actual AI driver that can scale.”
Again, you are making what seems to be an absolutely unsupported assumption that the long term development costs, systems management costs and ongoing rollout costs will distill down to a $200 difference in hardware.
This is essentially the point I have making all along. Along with the likelihood, Tesla can scale at a much cheaper cost.
I never bought into this concept and have never addressed it as a reason to own Tesla. It’s a liability nightmare both in terms of managing the safe upkeep of the vehicle for commercial use and who is responsible in the event of an accident. I’m still not entirely convinced that the robotaxi business itself will be anything other than a low margin commodity type business in the long run. I could see scenarios where Tesla completely exits operation and simply provides the system for others to implement.
You’re hung up on an argument almost nobody seems to be making here. It’s not that no additional sensors makes sense. It’s that Tesla has chosen to forgo LIDAR and laid out the reasons why with the default position that simplicity vs complexity is actually better in the long run. It is not a blanket opposition to multiplicity. In addition, they have yet to prove their approach works.
No, we’re not. We’re echoing Waymo’s argument that if your system only has one sensor input, the capacity of that system reaches a lower ceiling of performance than a system with multiple types of sensor input. It has nothing to do with the cost of the hardware.