Safety-critical AI is getting its own hour at TechCrunch Disrupt 2026. On the Real World AI Stage, Shield AI chief technology officer Nathan Michael, Waabi founder and CEO Raquel Urtasun and General Motors director of robotics strategy Mikell Taylor will share a panel called “Building AI Systems When Failure Is Not an Option”. TechCrunch announced the line-up on 24 September 2026, and the conference runs from 13 to 15 October at Moscone West in San Francisco.
The three speakers work on very different machines. Michael’s software flies military aircraft, Urtasun’s drives heavy trucks and soon robotaxis, and Taylor’s team builds robots for car factories. What they share is the question TechCrunch put at the top of its announcement: how do you know when an autonomous system is actually ready to leave the lab? A chatbot that gets something wrong produces a bad answer. A system controlling an aircraft, a vehicle or a robot can ground a plane, cause a crash or compromise a mission.
This article covers what the session will discuss and who the speakers are. It sets out the record each company brings to the stage and the evidence problem that makes safety-critical AI hard to sign off. It also covers the standards and safety culture behind a release decision, and the questions worth asking if you build or buy autonomous systems.
Table of contents
- What the Safety-Critical AI Panel at Disrupt 2026 Will Cover
- Why Safety-Critical AI Is a Different Engineering Problem
- Shield AI: Safety-Critical AI in the Air
- Waabi: Proving a Truck Is Ready to Drive Itself
- General Motors: Robots That Have to Work Beside People
- How Safety-Critical AI Earns a Release Decision
- Standards and Regulators Shaping Safety-Critical AI
- The Safety Culture Behind Safety-Critical AI
- Questions Founders Should Bring to the Safety-Critical AI Session
- What It Means for Businesses Deploying Safety-Critical AI
- Safety-Critical AI at Disrupt 2026: FAQ
- References
What the Safety-Critical AI Panel at Disrupt 2026 Will Cover
TechCrunch’s description of the session lists four themes: creating a safety culture, testing and validating systems, navigating regulatory hurdles, and earning trust. These are the parts of the work that rarely make it into a product demo. They decide whether a safety-critical AI programme ships or stalls.
The session and the stage
The panel sits on the Real World AI Stage, one of six industry stages at the event. TechCrunch says Disrupt 2026 will have more than 200 sessions across those stages, roundtables and breakouts. It expects more than 10,000 founders, investors, operators and technology leaders, over 250 speakers and more than 300 exhibiting startups. The organisers frame the panel as a comparison of safety-critical AI across defence, autonomous driving and industrial robotics, three fields where “almost ready” is not good enough.
Four themes on the agenda
The four themes map onto a real sequence. Safety culture comes first, because it decides what engineers are allowed to say when a test goes badly. Testing and validation produce the evidence. Regulation decides who has to see that evidence and in what form. Trust is the outcome: whether an air force, a freight carrier or a factory floor will accept the system. Safety-critical AI needs all four, and a gap in any one can stop a deployment that looks finished on paper.
Where it sits in the Disrupt programme
The panel is one of five sessions TechCrunch grouped together on 22 September as the AI safety track founders should not miss. Seen side by side, they show how widely the organisers are drawing the idea of safety this year.
| Session | Stage | Speakers | Focus |
|---|---|---|---|
| What Anthropic Sees When Enterprises Actually Deploy Claude | AI Stage | Cat de Jong (Anthropic) | Moving from pilots to production |
| The Agent Security Problem Nobody Is Talking About | AI Stage | Ric Smith (Okta), Gavriel Cohen (NanoCo) | Permissions and infrastructure for software that acts |
| Securing the AI Enterprise | AI Stage | Rudy Mitra (AWS), Katie Moussouris (Luta Security), Wendy Nather | Security, governance and observability |
| Building AI Systems When Failure Is Not an Option | Real World AI Stage | Nathan Michael, Raquel Urtasun, Mikell Taylor | Safety culture, validation, regulation, trust |
| Robots Are Waiting for Their ChatGPT Moment | Real World AI Stage | Nvidia Inception’s global head of physical AI | Data, simulation and foundation models for robots |
Three of the five are about software and enterprise data. The two on the Real World AI Stage are about machines. That split is the useful context for the Michael, Urtasun and Taylor panel. It is the only session in the group about safety-critical AI, where a failure is physical rather than digital.
Why Safety-Critical AI Is a Different Engineering Problem
Most AI products are judged on average quality. Safety-critical AI is judged on its worst day. That changes what “good enough” means, how evidence is gathered and who gets to sign off.
A wrong answer versus a wrong action
When a language model gets a fact wrong, a person usually reads the output before acting on it. There is a human between the error and its consequence. In safety-critical AI such as an autonomous aircraft, truck or mobile robot, the system senses, decides and acts in a loop measured in fractions of a second. Often no one is in a position to catch a bad decision before it becomes a movement.
That is why the panel’s framing starts with the physical world. The same underlying techniques, including perception built on computer vision, carry very different stakes once they steer a vehicle.
Rare failures are hard to measure
The failures that matter most in safety-critical AI are rare. A system can behave perfectly for millions of miles or flight hours and still fail on the one situation nobody tested. Rarity is good for the public but awkward for engineers, because it means ordinary testing produces very little evidence about the events you most need to understand. The next sections show how each company tries to close that gap, and why simulation sits at the centre of all three approaches.
The long tail
Engineers call the unusual situations the long tail: a pedestrian in a costume, a runway with odd markings, a pallet left half in an aisle. Safety-critical AI has to handle the tail gracefully, which usually means recognising uncertainty and falling back to a safe state rather than guessing. How safety-critical AI behaves when it does not know is as important as how well it performs when it does.
| Question | Chatbot or office software | Safety-critical AI |
|---|---|---|
| What does a failure look like? | A wrong or unhelpful answer | A crash, a lost aircraft or an injured worker |
| Who catches it? | Usually the person reading the output | Often no one in time |
| What is measured? | Average quality and user satisfaction | Worst-case behaviour and failure rates |
| How is it released? | Staged rollout and quick rollback | A documented safety case and a sign-off |
| Who else must agree? | Rarely anyone outside the company | Customers, insurers and often a regulator |
Shield AI: Safety-Critical AI in the Air
Shield AI brings the defence view to the panel. TechCrunch describes Nathan Michael as an AI leader developing autonomous systems “for environments where performance has to be matched by assurance”. That phrase is a good summary of safety-critical AI in a military setting, where a fault can cost an aircraft or a mission.
Nathan Michael’s background
Michael leads the development and deployment of Hivemind, Shield AI’s platform-agnostic mission autonomy software. His background spans AI, control, perception and multi-robot systems. Before Shield AI he spent years at Carnegie Mellon University’s Robotics Institute, where he directed the Resilient Intelligent Systems Lab. That academic grounding in multi-robot systems matters, because much of Shield AI’s recent safety-critical AI work involves several aircraft operating together.
What Hivemind does
Shield AI, founded in 2015 and based in San Diego, describes Hivemind as software that “assumes the role of a human pilot or operator”, letting uncrewed systems sense, decide and act. Unlike a traditional autopilot that follows a preplanned route, the company says Hivemind can reroute around no-fly zones, handle obstacles and respond to unexpected conditions. By March 2026 it had piloted 26 classes of vehicle, including F-16s, jet-powered drones, helicopters, drone boats and ground vehicles.
The Collaborative Combat Aircraft milestones
The year’s big test for Shield AI’s safety-critical AI came through the US Air Force’s Collaborative Combat Aircraft programme, which is developing uncrewed jets to fly alongside piloted fighters.
| Date (2026) | Milestone | Why it matters for validation |
|---|---|---|
| 13 February | Selected as a mission autonomy provider for the programme’s risk-reduction phase | Chosen after a competitive evaluation |
| 26 February | First flight test aboard Anduril’s YFQ-44A over the Mojave Desert | Completed all test points, including mid-mission updates |
| 26 March | $1.5 billion Series G at a $12.7 billion valuation, plus a deal to buy simulation firm Aechelon | Buys high-fidelity simulation for testing before live flight |
| 17 June | Production contract for programme mission autonomy | Multiple aircraft operating together under human supervision |
The valuation step shows how quickly investors have priced in that progress. TechCrunch noted that Shield AI had raised $240 million at a $5.3 billion valuation in March 2025, so the Series G price was up 140% in a year.
Simulation as a validation asset
The Aechelon purchase is the detail most relevant to safety-critical AI. Aechelon builds high-fidelity simulation, physics-based sensor models and synthetic environments that the US military uses to train pilots and test aircraft and autonomous systems before live flight. Shield AI chief executive Gary Steele said the deal would advance its “Hivemind Foundation Model for Defense, which is trained in simulation and continuously refined through real-world operations”. In other words, simulation is not a side project. It is where most of the evidence for flight-ready safety-critical AI is produced.
Waabi: Proving a Truck Is Ready to Drive Itself
Waabi brings the road view. TechCrunch put one of the session’s central questions squarely in Raquel Urtasun’s area: how do you decide when an AI system is ready to operate without a human behind the wheel?
Raquel Urtasun’s background
Urtasun has spent 25 years in AI and autonomous vehicles. Before founding Waabi she was chief scientist and head of research and development at Uber ATG, Uber’s former self-driving unit. She is a professor of computer science at the University of Toronto, co-founded the Vector Institute for AI and has published more than 200 AI papers. Few people on the Disrupt stage will have spent longer thinking about how to validate safety-critical AI on public roads.
Waabi World and the simulation-first bet
Testing and validation are central to Waabi’s approach. Its Waabi World simulator, as TechCrunch described it in January, is a closed-loop system that builds digital twins of the world from data, simulates sensors in real time, manufactures scenarios to stress-test the Waabi Driver and lets the driver learn from mistakes without human intervention. Urtasun argues this lets Waabi build with far less money and data than earlier self-driving companies. “We don’t need the gazillion humans to develop the technology and the large fleets that AV 1.0 needs,” she told TechCrunch.
The delayed driverless launch
Waabi is also a useful safety-critical AI case study because it has been open about waiting. In January 2026 TechCrunch reported that Waabi had run commercial pilots in Texas with a human driver in the front seat. It had planned to launch a fully driverless truck on public highways by the end of 2025, but that was pushed back to “sometime in the next few quarters”. Urtasun said the Waabi Driver was ready but the purpose-built Volvo trucks “still need to be fully validated before launch”.
By July the message had moved on without the launch happening. FreightWaves reported that the Waabi Driver, trained only on a Peterbilt 579, drove a Volvo VNL Autonomous truck on highways and surface streets with no retraining. Urtasun said the driverless launch would run on a fully validated vehicle platform from the manufacturer, and added: “Driverless is almost here.” For safety-critical AI, a company choosing to wait for vehicle validation is itself a data point about how release decisions get made.
From trucks to 25,000 robotaxis
The January round was Waabi’s move beyond freight. It raised $1 billion: an oversubscribed $750 million Series C co-led by Khosla Ventures and G2 Venture Partners, plus roughly $250 million in milestone-based capital from Uber. That money supports the deployment of 25,000 or more Waabi Driver robotaxis exclusively on Uber’s network, with no timeline given. Urtasun’s safety argument for robotaxis rests on the vehicle as much as the software. “We believe in vertically integrating with a fully redundant platform from the OEM,” she said. “That is how you really build safe and truly scalable technology.”
General Motors: Robots That Have to Work Beside People
General Motors brings the factory view, and with it the human side of safety-critical AI. Industrial robots have long been fenced off from people. The newer generation of safety-critical AI in factories is expected to share space with them.
Mikell Taylor and Proteus
Mikell Taylor has spent more than two decades building robots designed to do useful work. Her career started, as TechCrunch notes, with a robotic senior prom date, and has since covered autonomous underwater vehicles and industrial systems. Before GM she led the Amazon Robotics team that developed Proteus, Amazon’s first autonomous mobile robot. Proteus was notable precisely because it was designed to move around people rather than in caged-off zones.
GM’s Autonomous Robotics Center
Taylor now leads robotics strategy for GM’s Autonomous Robotics Center. GM said in October 2025 that the centre, in Warren, Michigan, with a sister lab in Mountain View, California, employs more than 100 roboticists, AI engineers and hardware specialists. They build systems trained on decades of GM production data, including telemetry, quality metrics and sensor feeds from thousands of robots. The centre is also developing software and manipulation parts for collaborative robots, or cobots, which GM said it was deploying in US assembly plants.
What makes a robot worthy
Taylor’s recent talks give a good preview of what she is likely to bring to Disrupt. At the Robotics Summit and Expo in Boston in May, her keynote was titled “What Makes a Robot Worthy?” The Robot Report summarised her argument: because robots operate in the physical world, they must prove they are worthy of adoption into places with very high bars for safety, performance, uptime and return on investment. She also warned against “pilot purgatory”, where promising robots never get past trials.
TechCrunch adds her point that user experience and adoption have to be designed in from the start, not bolted on at deployment.
How Safety-Critical AI Earns a Release Decision
The hardest question on the panel is the one TechCrunch highlighted for Urtasun: when is a system ready? Every safety-critical AI programme needs an answer that survives scrutiny from customers, insurers and regulators.
Why you cannot drive your way to safety
The best-known attempt to put numbers on this came from RAND in 2016. Researchers Nidhi Kalra and Susan Paddock calculated how many miles self-driving cars would need to drive to prove their safety statistically. They used the 2013 US human fatality rate of 1.09 deaths per 100 million miles. “Our results show that developers of this technology and third-party testers cannot drive their way to safety,” Kalra said. The chart below uses RAND’s own figures.
RAND worked out what those distances mean in time. A fleet of 100 vehicles driving 24 hours a day at 25 miles per hour would need about 12.5 years for the first figure, around 400 years for the second and about 500 years for the third. That is why every company on this panel leans so heavily on simulation. Road miles and flight hours still matter, but they cannot carry the whole statistical load for safety-critical AI.
A layered evidence case
In practice, teams building safety-critical AI assemble what engineers call a safety case: a structured argument, backed by evidence, that a system is acceptably safe for a defined use. Simulation supplies breadth, covering millions of scenarios including ones too dangerous to stage for real. Closed-course and test-range work checks that the simulator matches reality. Supervised operations, with a safety driver or a human operator, show behaviour in the real world at low risk. Only then does a limited unsupervised deployment begin, usually in a tightly defined area or mission.
| Evidence stage | What it proves | Panel example |
|---|---|---|
| Simulation at scale | Behaviour across rare and dangerous scenarios | Waabi World; Shield AI’s Aechelon purchase |
| Test range or closed course | The simulator matches the real vehicle | Hivemind’s YFQ-44A flight over the Mojave |
| Supervised operations | Real-world performance with a human backstop | Waabi’s Texas pilots with a driver on board |
| Platform validation | The vehicle and its redundancy are sound | Waabi waiting on the Volvo truck |
| Limited unsupervised use | Safety in a defined area or mission | Human-supervised multi-aircraft teams |
Knowing when you do not know
A release decision for safety-critical AI also depends on how the system handles uncertainty. Well-designed safety-critical AI should detect when it is outside the conditions it was validated for, which engineers call its operational design domain, and move to a minimal-risk state. For a truck that might mean pulling onto the hard shoulder. For an aircraft, holding a safe orbit or returning to base. For a factory robot, stopping. Validating that fallback behaviour is as important as validating normal operation.
Standards and Regulators Shaping Safety-Critical AI
“Navigating regulatory hurdles” is one of the four agenda items, and each speaker answers to a different rulebook. None of these frameworks certifies an AI model directly. Instead they ask companies to show a disciplined process and a convincing argument.
Defence
For the Pentagon, the key policy is DoD Directive 3000.09, “Autonomy in Weapon Systems”, updated in January 2023. It says autonomous and semi-autonomous weapon systems must be designed so that commanders and operators can “exercise appropriate levels of human judgment over the use of force”. Shield AI’s June contract language, which stresses multiple aircraft working together “under human supervision”, reflects that expectation.
Roads
For vehicles, the standards landscape is older and more detailed. ISO 26262 covers functional safety of road vehicles, and ISO 21448 covers hazards that arise without any fault, such as a perception system misreading a scene. UL 4600, the standard for evaluating autonomous products, asks for a full safety case covering the whole system. Its third edition, released in March 2023, added autonomous trucking, which is directly relevant to Waabi. In the US, the National Highway Traffic Safety Administration also requires crash reporting for automated driving systems.
Factories
Industrial robots fall under ISO 10218, and collaborative operation between people and robots has specific guidance on contact forces. For GM’s cobots, the practical test is whether a robot can share space with workers without the fences and guarding that traditional industrial robots needed.
| Domain | Framework | What it asks for |
|---|---|---|
| Defence | DoD Directive 3000.09 (2023) | Appropriate human judgement over the use of force; senior review |
| Road vehicles | ISO 26262 | Functional safety across the development lifecycle |
| Road vehicles | ISO 21448 | Safety when nothing is broken but perception falls short |
| Autonomous products, including trucks | UL 4600, 3rd edition (2023) | A structured safety case for the whole system |
| US roads | NHTSA crash reporting order | Reporting of crashes involving automated driving |
| Factories | ISO 10218 | Safety requirements for industrial robots and their integration |
The Safety Culture Behind Safety-Critical AI
TechCrunch lists creating a safety culture first among the session’s themes. It is the part of safety-critical AI that is hardest to see from outside. Standards describe processes. Culture decides whether people follow them when a deadline is close.
Permission to stop the line
In mature safety industries, anyone who sees a problem can stop work without being blamed. Car factories have long used this principle on the assembly line. For safety-critical AI teams it means an engineer can hold back a release when a test result looks wrong, even when investors are waiting. Waabi’s decision to delay its driverless truck until the vehicle is validated is the kind of choice that only happens in a culture that allows it.
Learning from near misses
Aviation became safer largely by collecting and studying near misses, not just accidents. Safety-critical AI teams can do the same with disengagements, operator takeovers and simulation failures. Every time a safety driver takes control, or a simulated run ends badly, there is a lesson about the long tail. A safety culture treats those events as valuable data rather than embarrassments to minimise.
Keeping humans meaningfully involved
All three companies keep people in the loop of their safety-critical AI in different ways. Shield AI’s aircraft work under human supervision, Waabi ran its pilots with drivers on board, and GM’s cobots are designed around workers. The challenge is making that human role meaningful rather than nominal. A supervisor asked to watch a highly reliable system for hours tends to lose attention, which is a known problem in automation research.
Questions Founders Should Bring to the Safety-Critical AI Session
TechCrunch pitches the panel at founders and technology leaders who build autonomous systems. If you are attending, or watching the recording later, these questions get to the heart of how safety-critical AI decisions are made.
How do you define “ready”?
Ask each speaker what evidence they need before removing a human. Is it a mileage or flight-hour target, a simulation coverage metric, a safety case signed off by an internal board, or a regulator’s approval? The answer shows whether readiness is a number, a process or a judgement.
How much do you trust your simulator?
Simulation carries most of the safety-critical AI evidence load for all three companies. The critical question is how they prove the simulator reflects reality, and what happens when real-world data disagrees with it.
What happens when the system is unsure?
Push on fallback behaviour. A safe response to uncertainty is often more important than peak performance in safety-critical AI.
How do you keep safety independent of commercial pressure?
Shield AI has defence contracts to meet, Waabi has a robotaxi deal with Uber and GM has production targets. Ask who can say no, and whether that person reports to the same executive who owns the launch date.
What It Means for Businesses Deploying Safety-Critical AI
Most organisations will never build an autonomous aircraft or truck. Many will buy systems that act in the physical world, from warehouse robots to automated inspection equipment. The lessons from this panel carry over directly to buyers of safety-critical AI.
Ask vendors for the safety case
A serious vendor of safety-critical AI should be able to explain the conditions their system was validated for, how it was tested and how it behaves outside those conditions. If a supplier can only show a demo video, treat that as a warning sign. Structured governance for these systems belongs in the same place as the rest of your risk management, alongside your approach to trust and security.
Plan for the human side
Taylor’s point about user experience applies to every workplace. Staff need training, clear rules on when to intervene and a way to report problems. The cleverest robot will fail commercially if the people around it do not trust it. Our coverage of Caterpillar’s approach to AI deployment shows how long that trust takes to build in heavy industry.
Follow the wider physical AI trend
The panel is part of a broader shift from AI in software to safety-critical AI in machines. For more background, see our explainer on physical AI and hardware that learns, our report on whole-body control for humanoid robots and our analysis of the Atoms robotaxi business. Our AI models and tools hub collects the latest releases.
Safety-Critical AI at Disrupt 2026: FAQ
When and where is the session?
“Building AI Systems When Failure Is Not an Option” is on the Real World AI Stage at TechCrunch Disrupt 2026, which runs from 13 to 15 October 2026 at Moscone West in San Francisco. TechCrunch had not published a session time in its 24 September announcement.
Who is speaking?
Nathan Michael, chief technology officer of Shield AI; Raquel Urtasun, founder and CEO of Waabi; and Mikell Taylor, director of robotics strategy at General Motors.
What is Hivemind?
Hivemind is Shield AI’s mission autonomy software. It takes the role of a pilot or operator on uncrewed aircraft and other vehicles, and it was selected in 2026 for the US Air Force Collaborative Combat Aircraft programme.
What is Waabi World?
Waabi World is Waabi’s closed-loop simulator. It builds digital twins from real data and generates scenarios to train and stress-test the Waabi Driver, the software that drives its trucks and planned robotaxis.
Why can’t autonomous vehicles just be road-tested?
Because serious failures are so rare. RAND calculated in 2016 that proving a self-driving car is 20% safer than human drivers would take about 11 billion miles, around 500 years for a 100-car fleet. That is why safety-critical AI relies on simulation and structured safety cases alongside real-world testing.
References
Five AI safety sessions every founder should have on their Disrupt 2026 agenda
Hivemind successfully completes first CCA flight test aboard Anduril’s YFQ-44A aircraft
Shield AI to acquire software simulation company Aechelon and raise $2B at $12.7B valuation
Defense startup Shield AI lands $12.7B valuation, up 140%, after US Air Force deal
Waabi raises $1B and expands into robotaxis with Uber
Waabi proves autonomous truck generalization with Volvo Autonomous Solutions
GM announces eyes-off driving, conversational AI, and unified software platform
Learn why robots need to earn trust from GM expert Mikell Taylor
UL 4600 Edition 3 Updates Incorporate Autonomous Trucking
Pentagon updates guidance for development, fielding and employment of autonomous weapon systems
More AI coverage: explore Progressive Robot's AI Models, Tools & Releases hub — hands-on reviews, setup guides and benchmarks in one place.