Most food safety investment goes into detection. The regulatory record shows that detection is not where sources are identified. Records are. That makes record design a process problem and it decides whether any monitoring investment returns anything at all.
In October 2024, an outbreak of E. coli O157:H7 was traced to yellow onions supplied to a national restaurant chain. FDA closed the investigation that December. Its conclusion repays close reading: the strain linked to illnesses was not found in onion or environmental samples. Epidemiologic and traceback evidence showed the recalled onions to be the likely source.
A second investigation ran the following month. FDA closed it in early 2025 with no public communication. Its internal incident summary, later obtained under a freedom of information request, is blunter than the public language. The outbreak strain was not detected in any inspectional samples. Romaine lettuce was nonetheless confirmed as the source, on epidemiologic and traceback data. A third produce investigation was open as this was written, with no confirmed positive product sample in it either.
Three investigations. No pathogen recovered from any food in any of them. In each case the source was named from records of what moved where and when.
The commodity changes. The pattern does not. The laboratory result is not what closes these investigations. It has not been for years.
Nor does the regulator drive the response. Across the 19 months to August 2026, every one of 2.4K food recall records in FDA's enforcement database was firm-initiated. Mandatory recall authority was not used once. The decision to recall sits with the manufacturer. So do the records that must support it.
The instinct after an event is to sample more. The arithmetic does not support it.
A field-validated 2022 study in Applied and Environmental Microbiology measured what a standard sampling plan actually detects. For a field carrying a single point source of contamination, a 60-grab sampling plan accepted the contaminated lot roughly 85% of the time. Reaching a 5% acceptance rate required 1.2K grabs. No commercial program samples at that density.
The more uncomfortable evidence comes from batches already known to be contaminated. In a recalled batch of powdered infant formula, researchers drew 415 samples of 333 grams each. 58% showed no detectable contamination at all. Contamination in dry product is clustered, so a clean result carries far less information than uniform-distribution arithmetic implies. That work dates from 2011 and the mathematics has not changed since.
Sampling tells an operator when they have been unlucky. It does not tell them whether they are in control.
Environmental monitoring is the standard answer and it is measured rarely enough that its record surprises people.
Researchers tracked eight small and medium dairy facilities across a full year, taking 2.1K environmental samples, of which 13% were Listeria positive. Over that year, only one of the eight facilities showed a statistically significant reduction in prevalence. Closely related isolates recurred over more than six months in five of the eight. That study was published in early 2024. It measures collection, not control.
What did improve outcomes was method. A 2026 study across all 50 state health departments found that a standard hypothesis-generating questionnaire was associated with a 155% higher Salmonella outbreak reporting rate. Routine interviews for all cases were associated with 85% and a standard cluster definition with 40%. Program participation and funded capability explained more of the variation than any individual surveillance practice.
Standardized method and funded capability outperformed instrumentation. That is an uncomfortable finding for anyone about to buy a platform.
Across our engagements in regulated manufacturing, the same shape appears before any technology question is reached. The record that would answer the regulator's question exists and it exists in people rather than in systems. On one regulatory program covering more than 120 markets, the operating rules had to be recovered through 26 stakeholder interviews and written down before any searchable, auditable system could exist at all.
A recent peer-reviewed review mapped the evidence for AI in postharvest food safety. Its governance criteria are more useful than its technology survey. Stronger evidence correlated with defined control points, external validation, auditable data flows and documented human oversight. The review asks three questions of any deployment: who may override an alert and how, who approves the intervention and how the correction is documented and whether an inspector can reconstruct after the fact which data and which logic produced the alert.
Its warning is the sentence to carry into any investment discussion. Without lifecycle documentation, an AI tool may improve initial detection but weaken long-term regulatory traceability.
The performance evidence justifies that caution. Across a systematic review of 28 pathogen-detection studies, only one used a fully independent test set from distinct sources. Eleven carried a high risk of bias and only two a low risk. Detection models tend to report high accuracy inside the facility that trained them and lose it elsewhere.
Two things should be said plainly. No system detects a pathogen on moving product at any commercially useful sensitivity, because validated methods need enrichment measured in tens of hours. And no case exists in the public record where predictive monitoring prevented a specific food safety event. Prevented events leave no record, so any vendor claim of one cannot be checked.
Recall size is not set by how clean a plant is. In one recent cascade, a contaminated ingredient from a single supplier produced Class I recalls at nine separate downstream manufacturers across two agencies. In another, glass traced to a supplied ingredient spanned 16 months of production. In both, volume was determined by how long the ingredient had been in use and how far it had traveled before anyone could trace it. Detection latency at the supplier interface is the multiplier.
That makes one question worth answering this week, without a consultant and without a system purchase.
How long to assemble a full lot genealogy for one case, without asking the person who knows?
The threshold is not ours to set. Major customers have already published theirs. A national warehouse club requires two independent trace exercises a year, each accounting for 100% of product within two hours. A large quick-service buyer requires location of 100% of any finished product within three hours. A global coffeehouse chain sets four hours, at 98-102% recovery. If the answer depends on one individual being reachable, it is not a two-hour process.
The regulator has published its own diagnosis. In the final rule, FDA states that a lack of consistent recordkeeping continues to hinder its traceback investigations, citing unstandardized records and an inability to link incoming product with outgoing product inside a single firm. In an August 2025 filing, FDA recorded that required data elements are not routinely maintained or shared across supply chains, that systems are not interoperable and that very few firms expected to meet the original January 2026 date.
What happened next is widely misread. The compliance date in the rule was never amended. FDA proposed an extension in August 2025 and never finalized it. What changed is that Congress barred FDA from spending appropriated funds on enforcement before July 20, 2028.
Customers did not wait. The largest US grocer has required traceability data on all food, not only the regulated list, since August 1, 2025. It publishes the consequence: it may hold or reject freight and levy penalties. A national supermarket chain set June 30, 2025. It states that because it requests the data for all foods, there are no exceptions. The regulatory date moved by roughly three years. The commercial date passed last year.
One detail makes the point better than any argument. FDA projects an 83% reduction in traceback time from the rule. It has never published the starting point that reduction is measured against. Even the regulator cannot verify its own improvement.
We do not begin with the monitoring layer. We begin by establishing whether the record exists, who owns it and how long it takes to produce under pressure. Measurement comes before design and design comes before any system is switched on.
That sequence is what both of our own AI deployments required. Rules held in individual memory across 26 stakeholders had to be written down before a platform covering more than 120 markets could be built. An automated compliance validation check caught a configuration error before go-live. The finding reached a named individual because that ownership was defined before the check was switched on.
AI is how we work. The result remains what we are accountable for.
If your last trace exercise ran longer than the standard your customers publish, the records assessment is the starting point before the next monitoring investment is approved.
Thirty minutes with a senior partner. No deck, no pitch.