Human drivers never pool their hard-won lessons; AV fleets could, if the law let it.
Incident facts are not trade secrets. Separate an incident into its five layers (event, conditions, behavior, failure mode, root cause) and it clarifies what to share and what to protect.
Here is a five-part regime (Report, Distribute, Act, Verify, Enforce) that aviation, pharma, and medical devices already run, and that Congress can start now.

A driver brakes hard as a child steps out from between two parked cars. He caught a flicker of motion in his peripheral vision, half a second of luck. The near-miss stays in his head. He learns from it.
But the next driver, in the next city, doesn’t catch the flicker. The child is hit. Then the next driver, the same mistake.
For 130 years, this is how we have learned to drive: one driver at a time. We have gotten dramatically better at surviving crashes. Seatbelts, crumple zones, airbags, and electronic stability control cut the death rate per mile roughly tenfold since the 1920s. We have barely gotten better at avoiding them, because avoidance does not transfer. Each driver’s hard-won lesson dies with the driver who learned it.
An autonomous vehicle is different. A lesson one vehicle learns can be pushed to every vehicle in its fleet overnight. The fleet’s competence only goes up. If we let it.
That asymmetry, human learning that resets with every driver against machine learning that compounds across a fleet, is the real reason to be optimistic about autonomous vehicles. And it is the potential the SELF DRIVE Act of 2026 has yet to capture.
The phrase to remember: an AV should only make the same mistake once.
Today’s letter digs into what the 2026 SELF DRIVE Act (H.R. 7390) gets right on data sharing, where it falls short, and how to fix it. We’ll cover the vocabulary that breaks the catch-all word “incident” into five precise layers a regulator can act on, the principle that should govern AV safety data (incident facts are not trade secrets, with one important caveat), the five mechanisms that turn one car’s near-miss into every fleet’s lesson, the governance choice that decides whether the regime works or becomes theater, and the precedents from aviation, pharma, and medical devices that prove all of this is buildable.
What the Bill Does, and What It Misses
Bipartisan AV legislation is hard work. Reps. Bob Latta and Debbie Dingell shipped the original SELF DRIVE Act in 2017; it passed the House unanimously and died in the Senate. Latta reintroduced it in 2021, and it stalled in committee. H.R. 7390 is the 2026 attempt, the most serious federal AV push since that first bill.
Section 30131 establishes a National Automated Vehicle Safety Data Repository. AV manufacturers must report every covered crash to the National Highway Traffic Safety Administration (NHTSA) within 30 days, or within 10 days of learning of it, whichever is later. That is a meaningful upgrade over the 2017 bill.
But the reporting requirement is incomplete. It covers crashes, not the near-misses and safety-relevant events that often precede them. AVs learn best from near-misses. The bill does not require them. A crash report records an outcome. It does not record the conditions that produced it, or the behavior that failed, which are the parts a fleet actually learns from.

Figure 1. The Reporting Gap — what the 2026 bill requires reporting, vs. 2017 and what is still missing

Figure 2. The Institutional Gap — investigation, CBI protections, and cross-operator dissemination
The SELF DRIVE Act inherits its DNA from the US auto industry, and it is not a flattering inheritance. The lesson of the industry’s worst safety failures is not that a reporting channel was missing. It is that the channel existed and did not work. The Takata airbag scandal, the largest automotive recall in U.S. history, escalated even though manufacturers using common inflators had early signals. The GM ignition-switch defect killed at least 124 people while the data sat visible in NHTSA’s Early Warning Reporting for years; GM later paid a $900 million criminal penalty for concealing it. In both cases the pipe was there. What failed was detection, mandatory action, and enforcement against concealment. A repository that only collects is not enough. NHTSA’s Standing General Order, the closest existing federal mechanism, has the same limit: the data goes to NHTSA and never across to other operators.
For an industry whose products kill roughly 40,000 Americans a year (39,345 in 2024, the first year under 40,000 since 2020), the safety-data infrastructure is primitive. And those deaths are not evenly shared. Pedestrian fatalities are the fastest-rising share, concentrated in lower-income neighborhoods and among Black and Native Americans walking on roads built only for cars. The data that explains them should not sit inside one company’s vault.
Waymo Earned Its Lead
Waymo has spent more than a decade and enormous capital on safety, and the payoff is measurable. Its 2025 peer-reviewed study, covering 56.7 million autonomous miles across four cities, reports roughly 85% fewer serious-injury crashes than human benchmarks. The confidence interval is wide (39% to 99%) and the comparison uses Waymo’s own constructed baseline, but it is a real result, and it holds within the operational design domains where Waymo actually drives: mapped, geofenced, mostly moderate-speed streets. That record is an achievement. The lessons behind it are what the rest of the industry needs.
Can the lessons from Waymo, Apollo Go, WeRide, and Tesla compound into a shared body of safety knowledge? Today, they don’t.
Here is the objection a Waymo executive will raise, and it is fair: why should the company that spent a decade earning a safety lead hand its hard-won lessons to fast followers? A well-designed regime answers it. Three features do the work.
What gets shared is the hazard, not the fix. The incident fact (the event, the conditions that triggered it, and the hazardous behavior) goes to the repository. The training data, the models, the failure mode inside the stack, and the fix all stay proprietary. Reporting that a hazard exists is the public good; diagnosing and solving it fastest remains each operator’s moat.
The regime is forward-only. Only incidents after the Act passes are reported. No one is forced to surface a decade of history.
The burden is fair by construction. Report volume scales with operation size and inversely with tech-stack quality: larger fleets contribute more because they operate more, weaker stacks contribute more because they have more to learn from. The safety externality, meanwhile, falls on people who never consented to it (a pedestrian, a cyclist, the driver in the next lane), which is exactly why pooling the hazard beats letting each operator hoard it.
What “Only Once” Actually Means: The Five Layers of an Incident
Pooling beats hoarding only if the pooled lesson actually transfers. And here the strongest objection lands: in an open world, no two incidents are identical. Differences in occlusion, timing, and behavior matter, and treating one logged event as a patch for the whole fleet would invite overfitting and new failure modes. Fleet learning is one layer of safety, on top of system design, validation, and redundancy, not a substitute for them.
The fix is to stop using one word for five different things. Pull an incident apart and it has five layers.
1. Event. The occurrence, timestamped and located. “A Waymo passed a stopped bus in Austin on January 12.” This is the raw report.
2. Triggering conditions. The situation that surfaced it. “Bus stopped, red lights flashing, students boarding, low sun, two-lane road.” This is exactly what ISO 21448 (SOTIF, the automotive “safety of the intended functionality” standard for hazards that arise without any component fault, just performance limits in hard conditions) calls triggering conditions. Technology-agnostic. A recurring set of them, generalized across incidents, is a scenario.
3. Hazardous behavior. The vehicle-level unsafe behavior, independent of cause. “Failed to stop for a stopped school bus.” This is ISO 26262 and UL 4600 territory, defined by what happened on the street. Generalized across incidents, it is a hazard class.
4. Failure mode. How this stack produced the behavior. Here, a remote-assist operator gave the car a wrong answer. Operator-specific, partly proprietary.
5. Root cause. The systemic reason beneath the failure mode. Not the wrong answer itself, but a protocol that let one operator override with no second check. This is what an investigator establishes, NTSB-style.
Now the slogan is precise. “The same mistake” means the same hazardous behavior in comparable triggering conditions, a hazard class, not an identical event and not the same failure mode. “Only once” means that once a hazard class is characterized, no other operator should meet it uninformed. It does not mean one incident patches the fleet.
Waymo’s school-bus case is the proof. In October 2025, NHTSA opened an investigation (PE25013) after Waymo vehicles drove past stopped school buses with red lights flashing and stop arms out; Waymo issued a voluntary software recall of 3,067 vehicles in December. Then in January 2026 it happened again, for a different reason: the NTSB found that a remote-assistance agent had wrongly told the vehicle the bus was not signaling. Same hazardous behavior, comparable triggering conditions, two different failure modes. To the public, the same mistake twice. That recurrence is the argument for shared hazard classes and an independent investigator, not against them. The hazard, “AVs misjudging stopped school buses,” should put every operator on notice, while each runs its own assessment of whether its stack could produce it.
And it reframes what the sharing is for: not a magic one-time lesson, but transparency, benchmarking, and safety oversight across the ecosystem, the conditions under which fleet learning can actually compound. The same layers also tell you exactly what a fleet should share and what stays locked inside it, which is the next question.
How the AV Industry Can Learn Together
The principle is simple, and the five layers make it operational. The two categories below are just a cut across them: the first three layers can be shared, the fourth is a trade secret, and the fifth is extracted by an investigator rather than reported.
Data sharing must distinguish two categories.
Category 1, incident facts. The first three layers, the event, the triggering conditions, and the hazardous behavior: what happened and the situation that surfaced it, not how anyone diagnosed or fixed it. This is what the bill must surface.
Category 2, trade secrets. The failure mode (layer four) and the internals that explain it: training-data corpora; perception, prediction, and planning models; routing algorithms and HD maps; hardware designs; pricing models; customer data. All of it genuinely deserves Confidential Business Information (CBI) protection. It stays proprietary to each operator.
That leaves the fifth layer, the root cause. It sits in neither category. No operator will publish its own, so an independent investigator digs it out and makes it public, which is the Verify mechanism below.
The line is real, but it is not clean, and a serious regime has to admit where it blurs. A triggering-condition report detailed enough to teach another fleet can edge into the operator’s failure mode: where a sensor faltered, where the operational design domain broke down. Aviation, pharma, and medical devices resolve this not by pretending the overlap away but by tiering disclosure: the hazardous behavior and its conditions are public; the causal engineering detail goes confidentially to the regulator and into de-identified aggregate trends; the full record is discoverable only by the regulator under confidentiality. The AV regime should copy that structure exactly. Exactly where the public tier ends, though, is the granularity question we return to below.
Within Category 1, the triggering conditions (layer two) carry most of the learning, and they fall into four families.
Environmental conditions: sun glare at a low angle, snow occluding lane markings, wet diesel residue that cuts tire grip.
Infrastructure conventions: construction-zone signage variations, hand-signal patterns from traffic officers, edge cases in traffic-control devices.
Road-user behavior: pedestrian intent cues (a half-step toward the curb that signals stepping off), cyclist swerve precursors (a head turn before swerving around a pothole), dooring precursors (a silhouette moving inside a parked car), regional driving cultures (the unspoken rules of a left turn in Boston).
Other safety-relevant phenomena: cybersecurity events (sensor spoofing, data-integrity attacks), sensor degradation, GPS denial, communication failures.

Figure 3. Anatomy of an Incident — the five layers, the two categories, and the four families
Do AV Companies Compete on Safety?
In aviation, companies don’t compete on safety. Pharma doesn’t compete on adverse-event rates. Medical devices don’t compete on MAUDE filings. In each of those industries, safety is a regulatory floor, not a marketing dimension.
AV companies do compete on safety. That is the market reality, and it is the whole problem. The companies that lead on safety treat their lessons as proprietary, because doing otherwise would forfeit the lead they bought with a decade and a balance sheet. When the law does not mandate sharing, sharing is unilateral disarmament.
Auto-ISAC shows that voluntary sharing can work in narrow scope: cybersecurity, where competitors face a common adversary. Safety has no common adversary. Unprotected voluntary sharing does not work in a market where each operator’s safety record is its differentiator. The end state has to be mandatory, and fair to every player. The path there can start with a voluntary, immunity-protected channel, the kind aviation runs in ASRS alongside its mandatory reports. We come back to how at the end.
Five Provisions: Report, Distribute, Act, Verify, Enforce
Industry-wide learning needs five mechanisms. None of them is exotic. All have working precedent in adjacent industries, and together they cover the five layers: sharing the first three, ring-fencing the failure mode inside each operator, and extracting the root cause independently.

Figure 4. The Mechanism–Layer Map — how the five mechanisms act on the five layers
One, comprehensive standardized reporting. The statutory definition of “safety-critical incident” is really the definition of the hazard classes, and the standardized schema is the controlled vocabulary that records the event, its triggering conditions, and the hazardous behavior in the same fields every time. Crashes are obvious; the harder definitions cover near-misses, anomalies, and operational-design-domain breaches, and aviation has decades of definitions to borrow. One caution, learned the hard way: do not turn this into a California-style disengagement scoreboard. California’s mandatory per-mile disengagement reports created a perverse incentive, rewarding the operator who tells a safety driver to keep hands off in the very moment the system is least sure, and even California now appears to be retiring the metric. The goal is a confidential, non-punitive record of safety-relevant events, not a gameable proxy that punishes the honest reporter. A rough analogue for the schema is the International Civil Aviation Organization’s (ICAO) Annex 13 accident-report taxonomy. Without standardization, the repository is a pile. Pair it with a confidential near-miss channel modeled on NASA’s Aviation Safety Reporting System: anonymous, immunity-protected, capturing what mandatory crash reporting cannot, including the deliberate avoidance maneuvers that never become crashes.
Two, mandatory distribution. NHTSA distributes the hazard class and its triggering conditions, anonymized, to every operator of a comparable system, defined by SAE automation level and overlapping operational design domain. Timelines tier by severity: days for fatalities and serious-injury incidents, thirty days for standard ones (the distribution clock, distinct from the operator’s reporting deadline above). Every operator, every time. One caution: the word “anonymized” hides a hard problem. These reports contain sensor footage of people who never consented to be recorded, a pedestrian’s gait, a location, a time, behavior near a school. The regime needs a public data-governance standard for de-identifying bystanders, not just a trade-secret carve-out for operators. The public gets aggregate disclosure; operators get the detail. Models exist: NTSB findings, MedWatch, MAUDE.
Three, required response, the Comparable Risk Assessment. Each operator maps the shared hazard class onto its own stack: could our failure modes produce that hazardous behavior in those triggering conditions, and what have we done about it? Operators have sixty days to file, publicly, on a schedule. This needs statutory immunity for the disclosure itself: operators that file in good faith should not face increased litigation exposure for the act of disclosing, though they remain fully liable for the underlying conduct. Without that immunity, every operator will rationally claim every issue is “non-comparable” to avoid building a litigation roadmap, and the regime collapses into formalism.
Four, independent verification. Investigation that establishes the root cause, the layer no operator will publish about itself, with public findings of probable cause. The body doing it must be neutral, public, and free of any stake in the outcome.
Five, enforcement authority. Verify without enforcement is theater. The authority must be able to compel compliance through fines, recalls, and suspension of operating authority. Who holds that authority, and how it relates to verification, is the governance question.
Governance: Who Runs the Regime
All five mechanisms need governance: authority over the rules, accountability for results, independence from operators. Who writes the schema and anonymizes the reports? Who decides comparability across operators? Who reviews the risk assessments? Who investigates incidents and carries the enforcement teeth?
Two institutional designs have precedent.
The combined model (FDA). One agency does Verify, Act, and Enforce. Pharma and medical devices use it. Single chain of command, faster mandates, but the investigator has a stake in the outcome.
The separated model (aviation). Three institutions, three roles. The FAA distributes, mandates action, and enforces; the NTSB verifies with no enforcement role and no stake; NASA’s ASRS runs confidential near-miss reporting, deliberately walled off from the regulator so reporters keep their anonymity. Slower than the combined model, but investigative neutrality and reporter protection are structurally guaranteed.
For AVs, the separated model is the right call, and not only on principle. It is the politically buyable one. An operator will not feed candid near-miss data into a system if the same agency can turn that data into an enforcement action against it. Walling the investigator off from the regulator is what makes honest reporting rational. The combined model is cleaner on an org chart and dies on contact with industry counsel. Beyond the institutional split, the design space includes appointments, accountability mechanisms, timelines, and cost-benefit calibration, the subject of a companion piece.
Precedents: Aviation, Pharma, Medical Devices
The closest analogue is commercial aviation, which went from a deadly novelty in the 1920s to the safest mode of transportation in human history. The transformation was driven by mandatory, structured, cross-operator learning. Aviation runs all five mechanisms and binds every operator to all of them.
Report: NASA’s ASRS now collects over 100,000 confidential near-miss reports a year; the FAA mandates incident reporting from operators.
Distribute: FAA Airworthiness Directives bind every operator of the affected aircraft type, regardless of whose plane had the problem; NTSB findings are public.
Act: operators comply with Airworthiness Directives by stated deadlines, modifying aircraft, retraining crews, grounding fleets when required.
Verify: independent NTSB investigations, with subpoena power and public findings of probable cause.
Enforce: the FAA can revoke pilot certificates, ground aircraft, fine operators, and suspend operating authority.
Aviation also has two reinforcing features the AV industry lacks. The first is shared safety technology: the Traffic Collision Avoidance System (TCAS) warns pilots of nearby aircraft, every commercial airliner carries it, and no airline competes on collision avoidance. The second is a cultural norm of shared responsibility: pilots, airlines, manufacturers, regulators, and investigators treat zero incidents as a collective goal, and treat incidents as learning opportunities for the whole industry rather than embarrassments to hide. The norm makes the regime collaborative on top of the mandate. The AV industry does not yet have it.
Pharma runs the same logic through the FDA’s MedWatch. Medical devices have MAUDE. All three regimes treat incident facts as fundamentally not-trade-secret. And aviation built one more thing the AV field still lacks: a shared dictionary of what counts as an incident, ICAO’s ADREP taxonomy, maintained and revised for decades. The 2026 SELF DRIVE Act has none of it.
What We Still Don’t Know: The Taxonomy Is the Hard Part
Naming the regime is not the same as writing its dictionary, and the dictionary is the hard part. Four questions are genuinely unsettled, and none is answered by this letter.
First, granularity. How specific should a hazard class be? “Failed to yield to a pedestrian” is too coarse to act on; “failed to yield to a pedestrian in a red coat at dusk on a four-lane arterial” is too narrow to share. Where the useful resolution sits is an empirical question, not a drafting choice.
Second, custody. Who maintains the controlled vocabulary, and revises it as the technology moves? Aviation took decades to build ICAO’s ADREP taxonomy, with a standing body (the CAST/ICAO Common Taxonomy Team) to keep it current. The AV field has no equivalent and no obvious home for one.
Third, reuse. The field does not start from zero. ISO 21448 (SOTIF) already names “triggering conditions”; SAE J3016 defines the operational design domain; UL 4600 structures the safety case; ASAM’s OpenSCENARIO and OpenODD and the PEGASUS project offer scenario formats. The work is to assemble these into a shared hazard-and-scenario taxonomy, not to reinvent them.
Fourth, gaming. Any definition a regulator mandates, operators will optimize to. The defense is to anchor the reportable unit in observable behavior and outcomes, the layers hardest to bend, and to let the independent verifier audit the coding. A taxonomy that can be lawyered is a taxonomy that will be.
This is a multi-year standards effort, the kind of work that sits between NHTSA, the standards bodies, the operators, and the universities. It should start now, in parallel with the bill, not after it.
The Stakes: Who Acts, and on Monday
The honest answer starts with what won’t work. Dropping a mandatory near-miss disclosure requirement into the next markup will not survive committee. The prior bills died over other things, federal preemption, FMVSS exemption caps, forced arbitration, the commercial-truck carve-out, not over data sharing. But a hard disclosure mandate is exactly the kind of provision industry would fight hardest if it were added now. The trade-secret and litigation fears are real. Pretending otherwise is how good provisions die.
The realistic path borrows aviation’s actual structure: a voluntary, confidential, immunity-protected channel running alongside a mandatory floor, not one converted into the other. NASA’s ASRS, walled off from the regulator, has captured near-misses that way for fifty years while separate mandatory reports handle the rest. A staffer on the House Energy and Commerce or Senate Commerce committee can build the voluntary half now by cross-referencing two existing authorities: the confidentiality regime that protects ASRS reporters, and the de-identified public-release mechanic of the FDA’s device-reporting system (the publicly searchable MAUDE database, 21 U.S.C. 360i). Submission is confidential, the de-identified fact is released, and the report cannot be used in enforcement against the reporter. Start with the safe harbor to build trust and volume; if uptake stays thin, layer in a mandatory floor, keeping the voluntary channel intact, exactly as aviation keeps ASRS alongside its mandates.
NHTSA does not need to build this from scratch, and it shouldn’t be asked to. It already runs the Standing General Order crash-reporting schema under existing authority. The lift is to extend that schema with de-identified near-miss fields and a confidentiality wrapper, not to stand up a greenfield system an under-resourced agency cannot staff.
And for the reader who runs a state DOT or a city’s curb, with no lever over a federal repository: the move is to align your own AV-pilot data agreements now. When you permit Waymo or Zoox to operate, write the permit so incident facts can be released to a de-identified public standard. Otherwise you will be the one holding an NDA that blocks the very sharing this regime depends on.
The AV industry has the same opportunity at a far earlier point in its lifecycle, and it can compress the industry-wide safety learning curve from decades into years.
Go back to the child between the parked cars. The promise of a shared learning regime is not faster product cycles or a cleaner org chart. It is that the next driver, in the next city, the one who would not have caught the flicker, is now riding in a vehicle that already knows the hazard, because another car, somewhere, met it first. An AV should only make the same mistake once. One car’s incident becomes every fleet’s lesson.
–Jinhua