Eighty-One Point Two: What a Restaurant Inspection Score Measures
In one state's data, restaurants that went on to cause a foodborne outbreak had a mean score of 81.2 on their last routine inspection, and that did not differ from restaurants with no reported outbreak. Mean scores awarded by individual inspectors ranged from 69 to 92. And yet posting the grades appears to work
Abstract
A study of one state's restaurant inspection data reports that mean scores of restaurants experiencing foodborne disease outbreaks did not differ from restaurants with no reported outbreaks, with the mean score of the last routine inspection before an outbreak at 81.2 across 49 outbreak-source restaurants. Mean scores of individual inspectors ranged from 69 to 92, none of the twelve most commonly cited violations were critical food safety hazards, and in 1,698 inspections scoring between 60 and 80 no critical violations were cited at all. Establishments scoring below 60 improved by a mean of 16 points on the next inspection while those above remained stable, which is the signature of regression to the mean. Against this, disclosure studies report benefits: point-of-service disclosure jurisdictions averaged 0.29 restaurant outbreaks per thousand establishments against 0.70 for online disclosure and 1.00 for none, and a New York City programme coincided with salmonella infections falling 14 per cent. A 2019 reanalysis contends the leading Los Angeles study's econometric specification was flawed and that it was confounded by a large salmonella outbreak preceding implementation, and grade inflation after adoption is documented. This article argues a score can be an effective incentive and a poor measurement simultaneously.
1. Introduction: a number with consequences
A restaurant's inspection score is posted, published, and in some jurisdictions displayed at the door. It affects custom, insurance, licensing and lease terms. It is treated as a measure of how safe the food is.
The finding this article is built around Mean scores of restaurants experiencing foodborne disease outbreaks did not differ from restaurants with no reported outbreaks.1
1.1 What this article argues
That the score is a poor measure of hazard, that disclosing it nevertheless appears to reduce outbreaks, and that both can be true because an incentive does not have to be accurate to work. Sections 14 and 23 are the case.
1.2 And why it is in this journal
Because §26 reports which violation category contributes most to the score.
2. The scale of the underlying problem
Establishing that this matters before questioning how it is measured.
National estimates put foodborne illness at approximately one in six Americans, 48 million people, each year, resulting in more than 128,000 hospitalizations and 3,000 deaths.6
And the older study's framing: more than 54 billion meals are served at 844,000 commercial food establishments annually, with 44% of adults eating at a restaurant on a typical day.1
2.1 The denominators are enormous
Fifty-four billion meals against a few thousand reported outbreaks means the base rate of a restaurant meal causing a recorded outbreak is extraordinarily low.1
Which is worth holding, because §5's finding is partly a statement about how hard it is to predict a rare event from a common measurement. That observation is ours.
2.2 So inspection is a serious undertaking
Nothing below suggests it should not happen. The question is what the resulting number represents.
3. Which our sources disagree about
On the restaurant share, by a wide margin.
One source states that nearly 75% of these cases can be traced to food prepared by restaurants, delis, or caterers.6
The earlier study reports that of a mean 550 outbreaks reported annually over a five-year period, >40% were attributed to commercial food establishments.1
3.1 Forty per cent against seventy-five
Different periods, different denominators, and possibly different definitions of what counts as traceable.16
We report both and note that a figure this central being uncertain by that margin is itself worth recording.
4. The central negative finding
The study, and its design is straightforward.
Inspection data were available from 49 restaurants that were identified as the source of foodborne disease outbreaks investigated by health departments in one state over four years.2
4.1 An outbreak-source restaurant is a strong endpoint
Not a complaint, not a suspicion, but an establishment a health department investigated and identified as the source. Whatever else is uncertain here, the outcome variable is about as solid as this field gets.2
4.2 And their scores were compared
Against restaurants with no reported outbreak, over the same period and the same inspection system.1
5. Eighty-one point two
The number.
The mean score of the last routine inspection before the reported outbreak was 81.2, and the mean score of the inspection previous to the most recent inspection was 81.6.2
5.1 Eighty-one is a respectable score
Comfortably above the threshold at which the same system triggers intervention, which §12 shows is around sixty.1
5.2 With one qualification about the direction
A score is deducted for problems, so higher is cleaner. Eighty-one means nineteen points of deductions were found in a premises about to cause an outbreak.2
5.3 And it is not distinguishable from anybody else's
Which is the finding in §1 restated as a figure.1
6. And the inspection before that
A detail worth pausing on.
Eighty-one point six, essentially identical to the eighty-one point two that followed.2
6.1 So there was no trajectory
These establishments were not deteriorating toward an outbreak in a way a series of scores would have revealed. They were stable, respectable, and about to make people ill.
That reading is ours and it closes off the obvious rescue, which would be that a single score is uninformative but a trend is not.
7. The inspector range
The second finding, and it is about the instrument rather than the restaurants.
Mean scores of individual inspectors were 69-92.1
7.1 Which is a reliability problem, not a bias problem
If every inspector were harsh by the same amount the scale would simply be shifted and comparisons would survive. A spread between assessors means two identical establishments receive different numbers depending on scheduling.
That distinction is ours and it is why §8 phrases the consequence as it does.
7.2 Averaged across all their inspections
So this is not one harsh day or one difficult establishment. It is the systematic tendency of different assessors applying the same instrument.1
8. What a twenty-three point spread means
Set against §5.
The gap between the most and least generous inspector is larger than the gap between a restaurant that caused an outbreak and one that did not.1
8.1 Which makes the assessor a bigger source of variation than the outcome
A score therefore carries more information about who performed the inspection than about whether the establishment will make somebody ill.
That is our formulation and it is the sharpest way we can state the measurement problem.
8.2 And a later analysis found the same thing
A study of grading systems in seventeen large jurisdictions, collecting data from ten, identified substantial inconsistencies in how the same establishment was scored over time.8
9. What the points are made of
The composition, which explains the rest.
Among inspections with scores of 60 to 80, a mean of 2.4 critical and 11.4 noncritical violations were cited; for inspections with a score <60, the means were 5.4 and 16, respectively.1
9.1 The system distinguishes the two categories itself
Critical and non-critical are the inspection instrument's own labels, so nothing in this section requires us to second-guess what counts as a hazard.1
9.2 Non-critical items outnumber critical ones by roughly five to one
In the middle band, and by three to one in the lowest.1
10. Seventeen hundred inspections, no critical violations
The figure that settles what the middle of the scale means.
In 1,698 inspections with a score of 60 to 80, no critical violations were cited.1
10.1 So a score of sixty-five can mean nothing hazardous was found
And the entire deduction came from items the same system classifies as non-critical.1
10.2 Which makes the scale non-monotonic in hazard
A restaurant with one critical violation and nothing else could outscore a restaurant with a dozen non-critical items and no critical ones, and the second is the safer place to eat.
That inference is ours and it follows directly from the figures.
11. And the common citations
Which confirms it from the other end.
None of the 12 most commonly cited violations were critical food safety hazards.1
11.1 Which follows from how inspection works
An inspector sees a moment. Critical hazards are usually processes rather than objects: a holding temperature over hours, a hand washed or not, a cooling curve overnight. A cracked tile is present whenever anybody looks.
So the instrument systematically over-samples the observable and under-samples the temporal, and that explanation is ours.
11.2 The twelve things inspectors write down most often
Are, by the system's own classification, not the things that cause illness.1
12. The regression artefact
A third finding, which would ordinarily be read as success.
Establishments scoring <60 had a mean improvement of 16 points on the subsequent routine inspection, with an additional mean increase of 5 on the next inspection.1
While restaurants with a score >60 tended to have fairly stable scores on subsequent inspections, with a mean drop of 2 points.1
12.1 The natural reading is that enforcement works
A failing establishment is inspected, warned, and improves by sixteen points.1
12.2 The alternative reading is arithmetic
If a score is a true value plus a large error term, and §7 establishes the error is large, then establishments selected for having scored unusually low will on average score higher next time whether or not anything changed.
That is regression to the mean, and it is our reading rather than the authors'.
13. Why that matters
Because the two readings prescribe different things.
If enforcement produced the sixteen points, the system is working and should be intensified. If regression produced them, the improvement is an artefact and the resources spent on re-inspection bought less than they appear to have.
13.1 The study cannot distinguish them
Separating them requires a comparison group of low-scoring establishments not subjected to follow-up, which no jurisdiction would run.
13.2 The test would not be difficult in principle
Compare establishments scoring just below a follow-up threshold with those scoring just above it, where the difference between them is largely noise. Our articles on evaluation describe that design, and inspection data is administrative data that already exists.
13.3 And this journal keeps meeting it
Our articles on treatment evaluation describe the same problem: an intervention applied to the worst cases will appear to work even when it does nothing, because the worst cases were partly unlucky.
14. So the score is a tidiness index
Our conclusion from §§5 to 11, stated plainly.
It measures how many things an inspector noticed that could be written down, most of which the system itself classifies as not causing illness, with a large component attributable to which inspector attended.
14.1 And the sources describe an improvement in what it measures
One evaluation notes that decreases in violations for inadequate hand-washing facilities and worker hygiene and improper storage or use of equipment or utensils are also likely to decrease risk.7
Hand-washing facilities are a physical item standing in for a behaviour, which is the closest an observable instrument gets to a process hazard.
14.2 That is not nothing
General order in a kitchen plausibly correlates with the practices that matter, and a place with sixteen non-critical violations is not being run carefully.
14.3 But it is not what the number is presented as
A posted grade is read by the public as a statement about whether the food will make them ill, and §5 says it is not that.
15. But disclosure appears to work
And this is where the subject becomes interesting.
An analysis of one county's programme showed that posting hygiene grades caused health inspection scores to rise, consumer demand to respond to changes in posted hygiene grades, and the number of foodborne illness hospitalizations to decrease.3
15.1 Three separate effects
On the operators, on the customers, and on the health outcome.3
15.2 And an economic mechanism
That premises posting poor scores suffer negative economic effects as consumers choose to eat elsewhere, and that this loss of revenue following a poor hygiene rating may motivate proprietors to improve their hygiene practices.3
16. The gradient
The most recent and broadest evidence.
Agencies that disclosed inspection results at the POS reported fewer outbreaks (mean = 0.29 outbreaks per 1,000 establishments) than those that disclosed results online (0.7) or not at all (1.0).4
And having any grading method for inspections was associated with fewer reported outbreaks than having no grading method, with letter grades lowest of all.4
16.1 A dose-response shape
More visible disclosure, fewer outbreaks, monotonically across three levels.4
16.2 And the mechanism is plausible before any data
Section 15.2's chain runs from a visible grade to lost custom to changed practice, and every step of it is ordinary consumer economics rather than anything about food safety.3
16.3 Drawn from national surveillance
Covering 2,608 single-state foodborne illness outbreaks over three years, with 1,638 attributed to food prepared in a restaurant setting.5
17. The New York figure
A third line.
A study reported salmonella infections falling by 14% during the first full year of letter grading, hitting the lowest level in two decades, while in the neighboring states of Connecticut and New Jersey, salmonella ratings were unchanged.3
17.1 The neighbouring states are the control
Which is the right design: a comparison population exposed to the same season, the same food supply and the same reporting system, without the intervention.3
17.2 And fourteen per cent is a substantial effect
For a single named pathogen, in a city of that size, against an unchanged comparison region.3
Which is the strongest individual result in this literature, and §21 is the reason we do not treat it as settled.
17.3 And the programme evaluation found behavioural change
With evaluations in three cities finding their disclosure programs were used by consumers and led to improved restaurant sanitary practices.7
18. And then the reanalysis
Which undoes a substantial part of it.
A 2019 reexamination of the county study claims serious analytical shortcomings; it contends that the econometric specification was flawed, which resulted in inaccurate findings.6
18.1 That is a strong charge
Not that the effect was smaller but that the specification producing it was wrong.6
18.2 Which is how a literature is supposed to work
An influential result reexamined, its methods challenged in print, and the challenge itself published in a journal. Our articles frequently complain that nobody checks; here somebody did.
18.3 And it concerns the most cited study in the field
Which appears in the reference lists of nearly every source we read on this subject.37
19. The specific confounder
Which is unusually concrete.
A policy document states that the leading study of Los Angeles was confounded by the state's largest salmonella outbreak in Southern California occurring before Los Angeles implemented restaurant grading.8
19.1 The mechanism is regression again
A baseline measured during or just after an exceptional outbreak is an unusually high baseline, and illness would have fallen afterwards whatever happened.
19.2 Which is §12 at jurisdiction scale
The same artefact, applied to the evaluation of the policy rather than to the evaluation of an individual restaurant.
That connection is ours and it is the reason we gave §12 its own sections.
19.3 And the same document reports the state of the evidence
That there is mixed evidence about whether food safety inspection scores are correlated with foodborne illness outbreaks, citing three.8
20. Grade inflation
A fourth problem, documented in the systems' behaviour.
The same multi-jurisdiction analysis found substantial evidence of grade inflation after the adoption of grading.8
And a paper devoted to it examines repeated interactions between inspectors and restaurateurs as the mechanism.6
20.1 Which is the predictable consequence of making the score matter
Once a grade carries commercial weight, the inspector is no longer recording an observation. They are imposing a penalty, on somebody they will see again, and the incentives on them change accordingly.
That reading is ours.
20.2 Which our own journal has an interest in
Because pest findings are the largest contributor per §26, and an inspector reluctant to impose a commercially damaging grade has an obvious category in which to be lenient.
If grade inflation is concentrated where the points are, it is concentrated in ours. That inference is ours and we found nobody testing it.
20.3 And it erodes the measure as the incentive strengthens
Which is an unusually direct trade-off: the property that makes disclosure effective is the same property that degrades what is being disclosed.
21. And the detection problem
Which the disclosure study raises against itself.
It found a positive correlation between the number of complaints received per 1,000 licensed restaurants and the number of restaurant outbreaks reported, meaning jurisdictions with higher rates of foodborne illness complaints also had higher rates of outbreaks linked to restaurants.5
And the authors' inference: that the ability to receive and investigate foodborne illness complaints may be an important predictor of the ability of an agency to detect foodborne illness outbreaks.5
21.1 With the strength of association reported
A correlation coefficient of 0.29 overall, stronger for bacterial toxin-mediated outbreaks (0.35) than for norovirus (0.10) or Salmonella (0.01).4
22. Which cuts in an awkward direction
Our observation, and the authors have opened the door to it.
If the number of outbreaks an agency reports depends on the agency's capacity to detect them, then fewer reported outbreaks can mean fewer outbreaks or less detection.5
22.1 And the outcome measure in §16 is reported outbreaks
So the gradient from 1.00 to 0.29 is a gradient in reporting as much as in occurrence.4
22.2 We are not claiming the effect is entirely artefact
There is no reason to expect point-of-service disclosure to reduce detection capacity, and it might plausibly increase it by raising public attention.
22.3 And the pathogen breakdown supports it
The complaint correlation was strongest for bacterial toxin-mediated outbreaks at 0.35 and essentially absent for one named pathogen at 0.01.4
Toxin-mediated illness has a short incubation and an obvious meal to blame, so it is the kind most likely to generate a complaint. The pattern is what a detection effect would look like, and that reading is ours.
22.4 What we are saying
Is that an outcome measure whose sensitivity varies by jurisdiction cannot support a between-jurisdiction comparison without adjustment, and that the paper identifies the problem without resolving it.5
23. The resolution we propose
Which reconciles §14 with §15.
A score can be an effective incentive and a poor measurement at the same time, because what makes it an incentive is that people care about it, not that it is accurate.
23.1 The disclosure evidence is about behaviour
Scores rise, demand responds, operators improve practices. All three are effects of the number being visible and costly.3
23.2 None of them requires the number to be valid
A restaurant improving to avoid a bad grade improves whatever the grade actually measures, and if the grade correlates loosely with real hazard, some of that improvement lands on things that matter.
23.3 The analogy is a school league table
A crude ranking that measures the measurable and drives real effort, some of which improves teaching and some of which improves the ranking. Both effects are well documented and neither requires the table to be a good measure of education.
That comparison is ours and we think it is the right one.
23.4 So the mechanism is indirect and still real
Which is the most useful thing we can say about this subject, and we found nobody stating it.
24. Why an inaccurate number can still work
Elaborating, because the claim is easy to misread.
An incentive operates on whatever it measures. If the score rewards visible tidiness, operators produce visible tidiness.
24.1 The question is whether the things it rewards overlap with the things that matter
Section 14.1 argues they partly do: a well-ordered kitchen is more likely to be one where temperature control and hand washing also happen.
24.2 And one specific overlap is documented
That having a certified kitchen manager on site is associated with fewer critical violations on inspection and has been identified as an important factor for preventing foodborne outbreaks.7
Which is a variable appearing on both sides, and is the kind of thing that would make a loose proxy useful.
25. And what it cannot do
Three things the score should not be asked to do.
Tell a diner whether a particular meal is safe. Section 5.
Rank two establishments. Section 8 makes a difference of a few points uninterpretable.
And serve as an outcome measure for evaluating anything. Because §20 says it inflates once it is used that way.
25.1 That third one is the general lesson
A measure used to judge people stops measuring. Our article on trap catch found counts shaped by the counting, our article on bait uptake found a proxy standing in for an outcome, and this is the version where the distortion is deliberate on both sides.
26. The finding that concerns us directly
And it is the most commercially relevant sentence in this journal.
From the New York programme evaluation: decreases in presence and severity of vermin violations contributed in large part to improvements in inspection scores, but vermin violations remain the largest average contributors to inspection score on initial inspection.7
26.1 The largest average contributor
Of every category of violation in the system, pest evidence takes the most points off a first-inspection score.7
26.2 Which follows from §11.1
Pest evidence is exactly the kind of thing an inspection is good at finding: physical, persistent, present whenever anybody looks, and requiring no judgement about a process that happened last night.
So an instrument biased toward the observable will weight it heavily, and that is our explanation for why the category dominates.
26.3 And reducing it drove the programme's improvement
In large part, per the same sentence.7
26.4 The authors' own conclusion
That this suggests a need for more restaurant operator education on this topic.7
27. What that means for a food service client
Honestly, which means saying both halves.
Pest control is the largest single lever on the score. If a client wants the number to improve, this is where the points are.7
And the score is a contested measure of safety. Everything from §4 to §14.
27.0 Three things a client should know
The score moves. Section 7 means two inspections of an unchanged kitchen can differ by more than the client's improvement efforts.1
A rebound after a bad score is expected. Section 12, and it should not be read as proof that whatever was done in between worked.
And a good score is not a clean bill. Section 5 is the whole answer.
27.1 Which is an uncomfortable sales position
We can truthfully tell a restaurant that our work moves their grade more than anything else, and we would be selling improvement in a number this article argues does not measure what people think it measures.
27.2 So the honest framing is separate
Pest control in a food premises is justified by contamination, by the allergen burden our threshold article quantified, by stock damage and by the regulatory requirement, none of which depends on the score being valid.
The grade improvement is a consequence rather than the reason.
27.3 And there is a genuine safety argument
Rodents and cockroaches are documented vectors of foodborne pathogens, which our articles on contamination describe. That is a hazard claim independent of the inspection system entirely, and it is the one we would make.
28. Our own position
The disclosure.
Commercial food service is valuable work and §26 gives us the strongest possible sales argument for it. Section 27.1 says why we should not lead with that.
28.1 And the grade-inflation point runs against us too
Section 20.2 suggests that if inspectors soften anywhere, they soften where the points are, which is our category. A client whose vermin score improved may have improved, or may have been treated more gently on a second visit.
We have no way to tell those apart on a single premises, and neither does the client.
28.2 And an article questioning the score is against our interest
A client who believes the grade is an accurate safety measure is a client who buys more readily. We are telling them it is contested, which is the position this journal has taken about our own guarantees, our own founding story and our own sanitation advice.
29. The Manitoba position
Brief, and mostly a gap.
Food premises inspection here is provincially administered and its scoring, disclosure practice and publication arrangements differ from the American systems every source in this article describes. We are not stating what they are.
29.1 What we could not find
Any Canadian validation study of inspection scores against outcomes, any Manitoba inspector variability analysis, and any evaluation of the disclosure arrangements used here.
29.2 One thing that does transfer
The physical argument in §27.3. Rodents and cockroaches carry pathogens into food premises regardless of which jurisdiction is scoring the premises, and our contamination articles describe the mechanism rather than the paperwork.
29.3 And whether the vermin finding transfers
Section 26 is one city's scoring system, and whether pest evidence is the largest score contributor under a different instrument is an empirical question nobody has answered locally.
30. Limitations and open questions
The central negative study is one state, twenty years ago. Forty-nine outbreak restaurants, one inspection system, one period.1
And forty-nine is a small number. A real difference of modest size could be present and undetected at that sample, so the finding is better read as no large difference than as no difference.2
And the systems have changed since. Risk-based inspection and critical-item weighting have been widely adopted, which may have addressed exactly the composition problem in §9.
That is the most important caveat here and we could not establish how far it applies.
The reanalysis reaches us secondhand. Through another paper's description of its claims and a policy document's summary, rather than from the paper itself.68
The regression reading in §12 is ours. The study reports the figures without proposing that explanation.1
Sources disagree on the restaurant share. Section 3.
And we did not examine the inspection instruments. What counts as critical, how points are allocated and how the scales differ between jurisdictions are the details that would determine how far any of this generalises.
Sections 1.1, 3.1, 6.1, 8, 10.2, 12.2, 13, 14, 19.1, 19.2, 20.1, 20.2, 22, 23, 24, 25 and 27 are our reasoning. The regression interpretation, the non-monotonicity argument, the detection critique of the disclosure gradient, and the incentive-against-measurement resolution are ours rather than sourced positions.
31. Conclusion
In one state's data, restaurants that went on to cause a foodborne outbreak scored 81.2 on their last routine inspection and 81.6 on the one before, and those figures did not differ from restaurants with no reported outbreak.12 Mean scores awarded by individual inspectors ranged from 69 to 92, which makes the assessor a larger source of variation than the outcome. None of the twelve most commonly cited violations was a critical food safety hazard, and in 1,698 inspections scoring between 60 and 80 no critical violation was cited at all, so the middle of the scale is constructed entirely from items the system itself says do not cause illness.1 The sixteen-point improvement by establishments scoring below 60 reads as enforcement working and has the exact signature of regression to the mean.
And yet the disclosure evidence is substantial. Point-of-service disclosure jurisdictions averaged 0.29 restaurant outbreaks per thousand establishments against 1.00 where results were not disclosed, a New York programme coincided with salmonella falling 14 per cent while neighbouring states were unchanged, and evaluations in three cities found consumers using the grades and operators improving practices.437 It is not clean: the flagship Los Angeles study has been reanalysed as flawed in specification and confounded by a very large salmonella outbreak preceding implementation, grade inflation follows adoption, and the disclosure study itself reports that agencies with more complaint capacity detect more outbreaks, which makes the outcome measure partly a measure of looking.685
What reconciles the two halves is that an incentive does not have to be accurate. Scores rise and operators improve because the number is visible and costly, not because it is true, and if it correlates loosely with real hazard then some of that improvement lands where it matters. The trade-off is that the property making disclosure effective is the same property that inflates the grade. For this trade the relevant sentence is narrower and sharper: vermin violations remain the largest average contributor to inspection score on initial inspection.7 Which means pest control is the biggest single lever anybody has on a number whose validity is contested, and the honest way to sell it is on contamination and the regulatory requirement rather than on the grade.
References
- Restaurant inspection scores and foodborne disease. Journal article in an emerging infectious diseases title, published version. Used for the context that more than 54 billion meals are served at 844,000 commercial food establishments annually in that country, that 46 per cent of food spending goes to restaurant meals, that 44 per cent of adults eat at a restaurant on a typical day, and that of a mean 550 foodborne disease outbreaks reported annually over a five-year period more than 40 per cent were attributed to commercial food establishments; for the finding that mean scores of individual inspectors were 69 to 92; for the finding that none of the twelve most commonly cited violations were critical food safety hazards; for the violation counts that among inspections scoring 60 to 80 a mean of 2.4 critical and 11.4 non-critical violations were cited while for scores below 60 the means were 5.4 and 16; for the finding that in 1,698 inspections with a score of 60 to 80 no critical violations were cited; for the subsequent-inspection figures that establishments scoring below 60 had a mean improvement of 16 points on the next routine inspection with an additional mean increase of 5 on the one after, while those above 60 had fairly stable scores with a mean drop of 2 points; and for the conclusion that mean scores of restaurants experiencing foodborne disease outbreaks did not differ from restaurants with no reported outbreaks and that the inspection system should be examined. https://wwwnc.cdc.gov/eid/article/10/4/pdfs/03-0343.pdf
- Article version of the above study hosted by the same agency. Used for the statement that inspection data were available from 49 restaurants identified as the source of foodborne disease outbreaks investigated by health departments in one state from 1999 to 2002; and for the finding that the mean score of the last routine inspection before the reported outbreak was 81.2 and the mean score of the inspection previous to that was 81.6. https://wwwnc.cdc.gov/eid/article/10/4/03-0343_article
- As clean as they look? Food hygiene inspection scores, microbiological contamination, and foodborne illness. Journal article in a food control title, read as extracts. Used for the report that an analysis of one county's programme showed posting hygiene grades caused inspection scores to rise, consumer demand to respond to changes in posted grades, and the number of foodborne illness hospitalisations to decrease; for the report that a study in one city showed foodborne illness down and restaurant revenue up since the start of letter grading, with salmonella infections falling 14 per cent during the first full year and hitting the lowest level in two decades while two neighbouring states were unchanged; for the statement that studies connecting hygiene ratings to foodborne illness are conflicting, naming an earlier county study that concluded routine inspections can predict outbreaks; and for the economic finding that premises posting poor hygiene scores suffer negative economic effects as consumers eat elsewhere, and that the loss of revenue may motivate proprietors to improve hygiene practices. https://www.sciencedirect.com/science/article/pii/S0956713518304432
- Foodborne outbreak rates associated with restaurant inspection grading and posting at the point of service: evaluation using national foodborne outbreak surveillance data. Journal article in a food protection title, read as abstract. Used for the study design assessing reproducibility of an earlier national survey using outbreak surveillance data for 2016 to 2018, with participating programmes accounting for approximately 23 per cent of single-state restaurant outbreaks; for the finding that agencies disclosing inspection results at the point of service reported fewer outbreaks at a mean of 0.29 per 1,000 establishments than those disclosing online at 0.70 or not at all at 1.00; for the finding that having any grading method was associated with fewer reported outbreaks than none, with letter grades lowest; and for the positive association between mean foodborne illness complaints per 1,000 establishments and mean restaurant outbreaks reported, with a stated correlation coefficient of 0.29 overall, stronger for bacterial toxin-mediated outbreaks at 0.35 than for norovirus at 0.10 or one named pathogen at 0.01. https://www.sciencedirect.com/science/article/pii/S0362028X22067011
- Trade magazine account of the above surveillance study. Trade source reporting on research, cited as attributed material. Used for the figures that during 2016 to 2018 there were 2,608 single-state foodborne illness outbreaks reported, with 1,638 attributed to food prepared in a restaurant setting, and that outbreaks in survey respondents' jurisdictions accounted for 23 per cent; for the finding that agencies using a grading system had lower mean numbers of restaurant outbreaks per 1,000 establishments with letter grades lowest; for the acknowledgement that the study had limited ability to distinguish among grading methods used; and for the finding of a positive correlation between complaints received per 1,000 licensed restaurants and restaurant outbreaks reported, with the authors' inference that the ability to receive and investigate complaints may be an important predictor of an agency's ability to detect outbreaks, and their emphasis on agencies having a complaint mechanism. https://www.food-safety.com/articles/8191-displaying-restaurant-inspection-grades-linked-to-fewer-foodborne-illness-outbreaks
- Grade inflation in restaurant hygiene inspections: repeated interactions between inspectors and restaurateurs. Journal article in a food policy title, read as abstract. Used for the account that a 2019 reexamination of the leading county study of hygiene inspection impact claims serious analytical shortcomings, contending that the econometric specification was flawed resulting in inaccurate findings; for the conclusion that administrative and design features of targeted transparency policies can affect the information provided to the public and potentially health outcomes; and for the national estimates that approximately one in six people, 48 million, suffer foodborne illness each year resulting in more than 128,000 hospitalisations and 3,000 deaths, with nearly 75 per cent of cases traceable to food prepared by restaurants, delis or caterers. https://www.sciencedirect.com/science/article/abs/pii/S0306919220301640
- Impact of a letter-grade program on restaurant sanitary conditions and diner behavior in one city. Journal article in a public health title. Used for the report that evaluations in two other named cities found their disclosure programmes were used by consumers and led to improved restaurant sanitary practices, that mandatory posting of grade cards in one county improved inspection scores after controlling for restaurant characteristics, and that both evaluations detected decreases in foodborne illness after implementation; for the statement that having a certified kitchen manager on site is associated with fewer critical violations on inspection and identified as an important factor for preventing foodborne outbreaks; for the observation that decreases in violations for inadequate hand-washing facilities and worker hygiene and improper storage or use of equipment are likely to decrease risk; for the finding that decreases in presence and severity of vermin violations contributed in large part to improvements in inspection scores but that vermin violations remain the largest average contributors to inspection score on initial inspection, suggesting a need for more restaurant operator education on this topic; and for the note that an increase in violation points related to food contact surface maintenance was likely an artefact of inspectors citing that violation under a miscellaneous section before grading. https://ajph.aphapublications.org/doi/full/10.2105/AJPH.2014.302404
- Regulatory policy document concerning a county's restaurant placarding programme, hosted by a university research centre. Policy document, cited as attributed material. Used for the statement that when the county's placarding efforts began there was limited evidence to inform regulatory design and implementation; for the report of a study examining grading systems in seventeen large jurisdictions and collecting inspection data from ten, which identified substantial inconsistencies in how the same establishment was scored over time and substantial evidence of grade inflation after the adoption of grading; for the statement that the leading study of one county was confounded by that state's largest salmonella outbreak occurring in the region before the county implemented restaurant grading; and for the statement that there is mixed evidence about whether food safety inspection scores are correlated with foodborne illness outbreaks, citing three works. https://dho.stanford.edu/wp-content/uploads/Elias_Ho.pdf
How to cite this article
APC Exterminators Research Division (2026). Eighty-One Point Two: What a Restaurant Inspection Score Measures. APC Review, Data, Statistics & Bioinformatics. Retrieved from https://apcexterminators.com/insights/restaurant-inspection-scores-validity-incentive-versus-measurement