First things first: what is physical AI?
Chatbots work with text. Physical AI is AI that has to act in the real world: self-driving cars, factory robots, humanoid robots, or systems that predict when equipment will fail. Instead of vast amounts of text from the internet, it works with data from sensors, cameras, test benches and telemetry. When it gets a call wrong, the cost can be a crash, downtime or damaged equipment.
100,000 hours of perfect data can leave a model blind
CoreWeave, a cloud company that supplies computing power for AI, launched its Physical AI Field Engineering service on September 10, 2026. The service sends engineers with automotive, aerospace and mechanical backgrounds directly into customer teams to build models on data those customers already own: test bench results, simulation output, production sensors and live telemetry. The team and its methods came from Monolith, the engineering AI company CoreWeave acquired in 2025.
Richard Ahlfeld, Monolith's founder and now SVP of Physical AI at CoreWeave, put it bluntly in a written interview with Superintelligence: the most common failure is "a training set that never contained the moments that matter." As he puts it, "If you want me to predict when your car fails and you give me 100,000 hours of flawless driving data, I am blind."
The logic is simple: models learn from the examples they have seen. If failures never appear in the data, the model has no idea what a failure looks like. No amount of normal data will teach it to recognize an anomaly.
NEURA Gym: a data loop that closes every day
He points to customer NEURA Robotics as an example. NEURA Gym runs roughly 100 cells where a robot attempts a customer's real task, and the loop closes daily: evaluate what happened, decide what data to collect, collect it, train, validate in simulation, and test on the physical robot again.
The conclusion from this process: volume of data matters less than finding the right data. The value of field engineers lies in deciding which edge cases (situations that rarely happen but cause real trouble when they do) deserve synthetic coverage, and which anomalies are signal rather than noise. Those are calls about the domain, not the tooling.

75 iterations for decisions made in under 30 seconds
Another example is the radio system built for the Aston Martin Aramco Formula One Team. Starting from seven hours of annotated audio, the team refined it over 75 iterations. According to Ahlfeld, those iterations were mostly about making the model survive the acoustics: engine noise, helmet mics, 300 km/h wind, multilingual drivers and team shorthand.
Every iteration was evaluated in the Weights & Biases Weave evaluation tool against word error rate and LLM-based quality scoring. The system was then run live at the Qatar and Abu Dhabi races before full deployment. A question that used to take an engineer several minutes now resolves inside a pit window that closes in under 30 seconds.
One caveat: the interview gives no error rate. "75 iterations" and "30 seconds" describe the process and the outcome, not accuracy.
Validate the simulation before trusting synthetic data
The critical moments in the real world are hard to wait for and hard to capture, so the industry often fills the gap with synthetic data from simulators or "world models" that can generate realistic scenes. But Ahlfeld's rule number one is: you have to extensively validate your simulations against controlled real-world data sets.
He sorts synthetic data into two categories, depending on whether it can be trusted:
What synthetic data can fill: gaps that physics models well, such as lighting, geometry and motion.
What still needs physical testing: granular, liquid and soft materials, erratic human behavior, and the differences between the hardware itself and its simulation.

He gives the example of a humanoid robot currently in training. Its original training corpus had only a handful of low-light episodes and none in direct sunlight, so the robot failed those tasks outright. The team augmented the corpus with NVIDIA's Cosmos world model to add episodes under non-ideal lighting, retrained it, and the robot now handles both conditions. He stresses that this worked because lighting is exactly what a world model gets right. If the failure had involved contact, material behavior or wear, they would have had to go back to real-world testing.
What it means for Taiwan

Taiwan's energy transition is in a phase of widespread demonstrations and test runs. From carbon capture and storage to hydrogen and new types of power generation equipment, projects are accumulating their first field data. Ahlfeld's experience offers three lessons:
Anomaly data is an asset, not noise. The anomalies, shutdowns and edge-of-envelope operating conditions seen during test runs are exactly the data future models will lack most. They should be recorded systematically, not filtered out to leave only "clean" steady-state data.
Validate the simulation before using it to fill data gaps. Use real-world measurements to establish where a simulation can be trusted. For known simulation weak spots such as subsurface fluid flow and material corrosion, field testing cannot be skipped.
AI recommendations need human oversight. Ahlfeld argues that every recommendation an AI makes about a physical system should be reviewed by a human and leave a traceable record. For power plants and demonstration sites, that is a baseline too.
Rather than chasing the biggest model, start by asking one question: does our data contain the moment that matters?
Sources
Superintelligence, "The most common failure isn't a bad model": CoreWeave's Richard Ahlfeld on physical AI (published October 4, 2026): https://read.getsuperintel.com/p/the-most-common-failure-isn-t-a-bad-model-coreweave-s-richard-ahlfeld-on-physical-ai
CoreWeave, CoreWeave Launches Physical AI Field Engineering (September 10, 2026): https://www.coreweave.com/news/coreweave-launches-physical-ai-field-engineering-to-turn-proprietary-data-into-production-ai
CoreWeave, AI Agents Decode F1 Radio in Near Real Time: https://www.coreweave.com/blog/how-ai-agents-on-coreweave-help-process-f1-radio-in-near-real-time
CoreWeave, CoreWeave's Stack for Physical AI: https://www.coreweave.com/blog/wayve-decart-neura-robotics-nissan-all-running-physical-ai-on-coreweave
NEURA Robotics, NEURA Gym: https://neura-robotics.com/neuragym/
