How Grounding Reduces Hallucinations at a Lower Cost

Max Leblond

15 Sep 2026

AI is confidently wrong about neighborhoods. We measured it, then fixed it.

A Local Logic technical whitepaper. Companion to the article “AI is confidently wrong about neighborhoods. We measured it, then fixed it.”

Study conducted July 2026. Statistics are 95% confidence intervals, cluster-bootstrapped over neighborhoods, and we quote the lower bound of the interval rather than the point estimate.


This is the full technical companion to our research on AI accuracy in real estate. For the plain-language overview, read AI is confidently wrong about neighborhoods.

Abstract

Large language models are now the front door to real estate. Buyers and professionals alike ask them what a neighborhood is like to live in, and the models answer fluently. On the location facts that decide a purchase, they are also frequently and confidently wrong. This paper reports a paired evaluation of seven language models answering the same 490 neighborhood questions twice, once from memory alone and once connected to Local Logic data through the Model Context Protocol (MCP), with every claim fact-checked against live data rather than a pre-written answer key.

Grounding a model in Local Logic data roughly doubled the number of verified facts per answer, cut contradicted claims by about 2 to 4 times per model, and roughly halved the share of answers with multiple serious errors. The largest gains went to the smallest, cheapest models, which is where accuracy is otherwise weakest. A dedicated live-web-search product scored below every grounded model on the same questions, which points to the conclusion the study is built to support: on location, the value is in the data, not in the retrieval.


Key takeaways

  • Grounding a model in Local Logic data roughly doubled verified facts per answer (about 1.8x to 2.3x).
  • It cut contradicted claims by about 2 to 4 times per model, and roughly halved answers with multiple serious errors (about 32% to 17%).
  • The cheapest models gained the most: a grounded GPT-5 nano reached 97.6% on the product categories, edging full GPT-5, at roughly one-seventh the cost per answer.
  • A dedicated live web-search product scored 82%, below every grounded model. On location, the value is in the data, not the retrieval.
  • Scale: 7 paired models, 490 questions, 46 neighborhoods, 7,350 verified answers, 153,530 scored claims.

Executive summary

For readers who want the result before the method.

  • The problem is measurable. Across the study, ungrounded models made 1,861 claims that live data flatly contradicts. These were not hedged guesses. They read as confident local expertise: a transit line that does not exist, a dog park with an invented name, a school that was never real.
  • Grounding roughly doubles verified facts. With Local Logic data connected, answers carried about 1.8 to 2.3 times as many true, checkable facts per answer. Five of seven paired models cleared the 1.5x bar set before the run.
  • Grounding sharply cuts hallucination. The claim-level contradicted rate fell for every model, by roughly 2 to 4 times. On the weakest model it fell from about 22.7% to 6.1%. The share of answers with two or more serious errors dropped by roughly half on average.
  • The cheapest models gain the most. Accuracy lift is inversely related to a model’s standalone capability. The weaker a model is on its own, the more the data helps, which makes grounding an economic lever, not just an accuracy one.
  • AI is worst where the market is. Ungrounded accuracy slides as neighborhoods get less famous. Grounded accuracy stays flat across famous, mid, and obscure neighborhoods, because the data is retrieved regardless of fame.
  • Cost is the headline. A small, inexpensive grounded model reaches accuracy comparable to a frontier model working from memory, at a fraction of the cost per answer. On the categories exposed in the product today, a grounded GPT-5 nano reached 97.6% and edged out full GPT-5’s own 96.9%, at roughly one-seventh the cost per answer.
  • Web search is not a shortcut. A dedicated live-search product scored 82% on the location-dependent categories, below all seven grounded models. Retrieval alone does not close the gap; the authoritative dataset does.
Figure 1. Verified-fact volume lift by model. Grounded verified-fact count divided by the same model
Figure 1. Verified-fact volume lift by model. Grounded verified-fact count divided by the same model’s ungrounded count; five of seven models clear the 1.5x bar.

Key results at a glance

Measure Result with grounding
Verified facts per answer About 1.8x to 2.3x more (up to 2.2x on the strongest models on the product categories)
Claim-level contradicted rate Down about 2x to 4x per model; one model 22.7% to 6.1%
Answers with 2+ serious errors Roughly halved on average (about 32% to 17%)
Accuracy, grounded models 89% to 97% claim accuracy on location-dependent categories
Cost Cheap grounded model rivals frontier-from-memory accuracy at a fraction of cost per answer
Scale 7 models, paired; 490 questions; 46 neighborhoods; 7,350 verified answers; 153,530 scored claims

1. Why grounding, and why now

Location has always been the part of a home decision that the listing cannot describe. Square footage and price are on every website. The commute, the walk to transit, what is a few minutes away, how a street actually lives at different hours: that is what people are choosing, and until recently the only way to get it was to ask a person who knew.

In 2026 a machine started answering that question at scale, for free, with total confidence. The catch is accuracy. A model that has read the whole internet still has no reliable, calibrated, comparable record of what is around a specific address, so it fills the gap the way models do, by generating a plausible answer. For casual browsing that is tolerable. For a decision measured in years and hundreds of thousands of dollars, a confident wrong answer is a liability, and neither the buyer nor the professional relaying the answer can tell which claims to trust.

Grounding is the technique that closes this gap. A grounded model answers from verified external data supplied at the moment of the question rather than from patterns in its training. It looks the fact up instead of recalling it. This paper measures what grounding on Local Logic data does to a model’s answers about neighborhoods, across the full range of model sizes and prices a team might actually deploy.

2. What we set out to test

The study was pre-registered before the run. The questions, in plain terms:

  1. Do ungrounded models invent facts on the categories only we cover? We expected material hallucination on livability, affordability, climate, and demographics, where the underlying facts are not reliably in any training corpus.
  2. Does grounding make answers both richer and more accurate? This was the co-primary question, measured as verified-fact volume and as claim accuracy.
  3. Do weaker and cheaper models benefit more? If grounding substitutes for missing internal knowledge, the lift should be largest where internal knowledge is thinnest.
  4. Does adding tools ever hurt on facts the model already handles? A no-harm check on public categories like transit, parks, schools, and points of interest.
  5. Can a cheap model plus data match a frontier model, for less money? The economic question underneath all of the above.

3. Study design and methodology

Paired A/B design. Each model answered the identical question twice: once from its training data alone (the ungrounded arm), and once with access to the Local Logic MCP tools (the grounded arm). Seven paired models span frontier to cheap: GPT-5, GPT-5 nano, Claude Opus, Claude Sonnet, Gemini Flash, GPT-4o mini, and DeepSeek Flash. Running the same model on the same question with tools on and off isolates the effect of the data from the effect of the model. The ungrounded arm used a direct prompt with no hedging instruction, which is deployment-realistic and was committed to before any results existed. Runs were single-shot, with no retries or self-consistency voting.

Benchmark. 490 questions across 46 neighborhoods in 39 cities, sampled along a continuous fame axis (log Wikipedia pageviews) so results are not dominated by famous places. Roughly 20% of neighborhoods are Canadian. Market and affordability questions are US-only by design, because that market data is US-only, and this is disclosed rather than papered over.

Scoring: claims fact-checked against live data, not an answer key. A frozen claim-verification judge (a model with access to the Local Logic MCP and web search, temperature 0, locked before the run) breaks each answer into atomic claims and checks each one against live data: our data for the categories we cover, web search to backfill the rest. This is deliberately stronger than grading against a pre-written gold answer, because the accuracy numbers then rest on whether each stated fact actually checks out, not on any model’s opinion. Scoring is tiered:

  • Deterministic categories (livability, climate, affordability, commute, demographics) are checked against the raw data values.
  • Entity categories (transit, parks, schools, points of interest) are verified against data.
  • Synthesis categories (character, comparative, “find somewhere similar”) are the only tier where the judge weighs open-ended reasoning.

Metrics. Three headline measures. Verified-fact volume is the count of supported claims per answer, reported as the grounded-to-ungrounded ratio; it is a coverage measure, not a correctness measure. Claim verification rate is supported divided by supported-plus-contradicted, reported as the accuracy percentage, with unverifiable claims reported separately and never counted in the denominator. Hallucination is the claim-level contradicted rate, with a secondary “materially wrong” measure counting answers with two or more central contradicted claims.

Statistics and targets. Every figure is a 95% confidence interval, cluster-bootstrapped over neighborhoods, because questions within a neighborhood are correlated and naive per-question intervals would be too tight. We quote the lower bound of the interval, so a claim strengthens rather than collapses under scrutiny. Pass and fail bars were fixed before the run: verified-fact volume lift at least 1.5x, deterministic accuracy at least 90%, synthesis accuracy at least 95%, and a no-harm margin within 0.7 points.

4. Results

4.1 Verified-fact volume

Grounded answers carried more true, checkable facts. Pooling across categories and comparing each grounded arm to its own ungrounded arm, five of seven models cleared the 1.5x bar on the lower bound of the interval.

Model Volume lift (CI lower bound) Verdict
DeepSeek Flash 2.30x (2.21x) Pass
Claude Sonnet 2.14x (2.06x) Pass
GPT-5 1.95x (1.89x) Pass
Claude Opus 1.85x (1.80x) Pass
GPT-5 nano 1.77x (1.70x) Pass
GPT-4o mini 1.51x (1.42x) Below bar
Gemini Flash 0.94x (0.89x) Below bar

The two that fall short are informative rather than disappointing. Both are among the terser models: they emit fewer claims overall, so a volume ratio can sit near or below 1.0 even as the accuracy of what they do say improves. Volume is a coverage measure, and it should be read alongside accuracy, not on its own.

4.2 Accuracy

Grounded models reached high claim accuracy on both the deterministic and synthesis tiers. Bars were 90% deterministic and 95% synthesis.

Grounded model Deterministic (CI-LB) Synthesis (CI-LB)
GPT-5 0.968 (0.955) 0.988 (0.981)
Claude Opus 0.944 (0.937) 0.973 (0.966)
Claude Sonnet 0.929 (0.906) 0.970 (0.953)
DeepSeek Flash 0.929 (0.916) 0.967 (0.961)
GPT-5 nano 0.909 (0.895) 0.962 (0.953)
Gemini Flash 0.902 (0.864) 0.931 (0.899)
GPT-4o mini 0.852 (0.813) 0.908 (0.871)

Most grounded models clear both bars. The two cheapest general-purpose models (Gemini Flash, GPT-4o mini) fall short of the strict synthesis bar on the lower bound, which is consistent with the pattern throughout: grounding lifts every model, and the weakest models, while improving the most in absolute terms, still trail the strongest.

4.3 Hallucination reduction

The claim-level contradicted rate on the location-dependent categories fell for every model when grounded.

Model Contradicted, from memory Contradicted, grounded Reduction
DeepSeek Flash 22.7% 6.1% 3.7x
GPT-5 nano 16.6% 7.8% 2.1x
GPT-4o mini 15.6% 13.2% 1.2x
Gemini Flash 12.8% 7.9% 1.6x
Claude Sonnet 12.7% 5.9% 2.1x
Claude Opus 9.5% 3.9% 2.5x
GPT-5 5.9% 2.2% 2.7x

Counted per answer rather than per claim, the share of answers with two or more serious errors roughly halved: about 32% of ungrounded answers were materially wrong, against about 17% grounded.

A note on how this is reported. One pre-registered gate for the hallucination measure required both a high absolute rate and at least a 3x reduction; that gate was retired before analysis because its 3x multiplier assumed grounded rates in the low single digits, which the data did not support. Grounding roughly halves hallucination, which is real and reported here as a measured reduction rather than as a threshold pass. We disclose the retired gate rather than drop it.

Figure 2. Claim-level contradicted rate, from memory versus grounded, per model. Every model improves; the biggest movers are the cheapest models.
Figure 2. Claim-level contradicted rate, from memory versus grounded, per model. Every model improves; the biggest movers are the cheapest models.
Figure 3. Share of answers that are materially wrong, from memory versus grounded. Grounding roughly halves answers carrying two or more serious errors.
Figure 3. Share of answers that are materially wrong, from memory versus grounded. Grounding roughly halves answers carrying two or more serious errors.

4.4 The cheapest models gain the most

Accuracy lift is inversely related to a model’s standalone capability. Using each model’s ungrounded accuracy as a capability proxy, the correlation between capability and accuracy lift is strongly negative (Pearson r of about -0.77). The weaker a model is on its own, the more grounding helps.

Model Ungrounded baseline Accuracy lift (points)
DeepSeek Flash 0.773 +16.6
GPT-5 nano 0.834 +8.8
Claude Sonnet 0.873 +6.8
Claude Opus 0.905 +5.6
Gemini Flash 0.872 +4.9
GPT-5 0.942 +3.7
GPT-4o mini 0.844 +2.4

The trend is directional rather than a passed statistical gate: with only seven models, the slope’s confidence interval still crosses zero. The descriptive signal is clear and consistent with the mechanism, so we report it as directional support, not as proof.

Figure 4. Accuracy lift versus model capability. The weaker a model is on its own, the more grounding helps (Pearson r about -0.77).
Figure 4. Accuracy lift versus model capability. The weaker a model is on its own, the more grounding helps (Pearson r about -0.77).

4.5 The obscurity cliff

Ungrounded accuracy depends on fame. Models know famous neighborhoods and guess on obscure ones, so accuracy from memory slides as a place gets less known. Grounded accuracy stays flat, because the data is retrieved regardless of fame.

Fame tier Accuracy from memory Accuracy grounded Volume lift
Famous 88.2% 94.8% 1.90x
Mid 87.3% 92.7% 1.88x
Obscure 83.7% 92.4% 1.93x

This matters commercially because volume is not concentrated in famous places. A large share of transactions happen in the long tail of lesser-known neighborhoods, which is exactly where ungrounded models are least reliable and where grounding’s edge is largest.

Figure 6. Accuracy by neighborhood fame tier. Grounded accuracy holds flat and high; ungrounded accuracy slides toward obscure neighborhoods.
Figure 5. Accuracy by neighborhood fame tier. Grounded accuracy holds flat and high; ungrounded accuracy slides toward obscure neighborhoods.
Figure 7. Per-model accuracy from famous to obscure. Every from-memory line slopes down; every grounded line holds.
Figure 6. Per-model accuracy from famous to obscure. Every from-memory line slopes down; every grounded line holds.

4.6 No harm on the categories models already handle

Grounding did not degrade answers on the public-fact categories that models already handle well. On the entity tier (transit, parks, schools, points of interest), the grounded-minus-ungrounded difference was positive for every model, from about +2 to +23 points, comfortably clearing the no-harm margin. Adding the tools helped even here, and never hurt. This is reported as a proxy for the pre-registered no-harm test, which used a separate public-fact set that was deferred from this run.

Figure 7. No harm on the categories models already handle. Grounded minus ungrounded claim rate on the entity tier (transit, parks, schools, points of interest,) with 95% intervals. Every model is positive, well clear of the -.7 point harm bound.

4.7 Cost: a cheap grounded model rivals a frontier one

This is the finding that reorders budgets. Grounding lifts every model, and it lifts the cheap ones so much that they approach frontier accuracy at a fraction of the price.

Grounded model Location-dependent accuracy (CI-LB) Cost per answer
GPT-5 0.978 (0.963) $0.535
Claude Opus 0.961 (0.954) $0.517
Claude Sonnet 0.941 (0.910) $0.074
DeepSeek Flash 0.939 (0.921) $0.010
GPT-5 nano 0.922 (0.906) $0.004
Gemini Flash 0.921 (0.879) $0.006
GPT-4o mini 0.868 (0.829) $0.004

Read against the best frontier-from-memory baseline (GPT-5 at 0.942 on the location-dependent categories), the strict non-inferiority test is cleared on the lower bound only by the two expensive grounded models. The cheaper grounded models land within a few points of that frontier baseline at one-third to one-twentieth of the cost per answer, a trade most teams running at volume will take.

On the categories exposed in the product today (setting aside climate and affordability, which are measured in the study but not yet placed in front of AI), the picture is sharper still: a grounded GPT-5 nano reached 97.6% and edged out full GPT-5’s own 96.9%, at roughly one-seventh the cost per answer. The takeaway is the same at both cuts of the data: you do not need the biggest model, you need the right data feeding it.

Accuracy and cost per answer for each model are shown in the table above.

4.8 Live web search is not a substitute

To answer the obvious objection, why not just let the model search the web, the study included a dedicated live-search product as a control. It scored 82% on the location-dependent categories, below every one of the seven grounded models (the weakest grounded model was 87%), at a cost per answer higher than most of the grounded cheap models. Retrieval by itself does not reach a calibrated, comparable location answer, because that answer is not sitting on a page to be found. The moat is the dataset, not the search.

5. What the failures actually look like

The aggregate rates are easier to feel through specific errors. Every example below came from an ungrounded model in the study and was contradicted by live data. They fall into three recognizable failure modes, and none of them hedged. Each read like local expertise, which is exactly why the error is dangerous: the person receiving it has no signal that it is wrong.

Fabrications: things that do not exist

  • Capitol Hill, Seattle, WA (Gemini Flash). Claimed the station is served by the Sound Transit “L Line.” There is no L Line; Capitol Hill is served by the 1 Line and 2 Line.
  • North Park, San Diego, CA (GPT-5 nano). Cited a “Balboa Park Station” on the San Diego Trolley as a transfer point for the Green and Blue Lines. No such station exists; the nearest is City College.
  • Highland, Denver, CO (GPT-4o mini). Named the closest light rail stops as “Highland Bridge” and “30th and Downing.” There is no RTD station called Highland Bridge; it is a pedestrian bridge.
  • Montrose, Houston, TX (DeepSeek V4 Flash). Called the dog park in Buffalo Bayou Park the “John T. Scott Dog Park.” It is the Johnny Steele Dog Park; the invented name does not exist.
  • Walnut Hills, Cincinnati, OH (GPT-4o mini). Placed a Kroger at 4647 Reading Rd near the neighborhood. There is no active Kroger there; the local store closed around 2017.
  • Jamaica Plain, Boston, MA (GPT-4o mini). Recommended the “Jamaica Plain Community Charter School” as a popular choice. No school by that name exists in the district lists, federal data, or anywhere on the web.

Confident miscalibration: real measures, wrong numbers

  • Adams Morgan, Washington, DC (Claude Sonnet). Said the peak car commute can stretch to 25 to 35 minutes or more. Local Logic routing at 8:30 a.m. returns 11 minutes, barely above the 10-minute off-peak time.
  • Port Richmond, Philadelphia, PA (GPT-5). Put the drive to Center City at about 15 to 25 minutes. Routing returns 12.
  • Central West End, St. Louis, MO (GPT-5 nano). Claimed a walk score in the 90s in the core. The neighborhood-level score is 78; a single address reached 88. No 90s.
  • NoDa, Charlotte, NC (GPT-5). Described a “vibrant score indicating a lively atmosphere.” The measured vibrant score is 1.73 out of 5, which reads as calm for most of the day.

Comparative and similarity errors

  • Williamsburg and Greenpoint, Brooklyn, NY (GPT-5). Called Greenpoint’s transit “more limited, mainly the G train.” Greenpoint scores 5.0 on transit, with multiple subway services, roughly ten bus lines, and ferry access.
  • Deep Ellum, Dallas, TX (GPT-5). Named East Austin as the closest match to Deep Ellum. On the measured scores the two diverge sharply (Deep Ellum nightlife 4.89 and quiet 1.28 against East Austin’s 3.50 and 4.02); Austin’s Red River Cultural District aligns far better.
  • Wynwood, Miami, FL (GPT-5). Said Wynwood has far fewer parks than Coconut Grove, Coral Gables, Miami Shores, or Morningside. Wynwood’s parks score of 3.2 is equal to or higher than all four.
  • Ohio City, Cleveland, OH (GPT-5). Called Clark-Fulton similar in feel to Ohio City. Their character scores differ across the board (quiet 4.62 against 2.98, cafes 2.00 against 3.15), and Clark-Fulton does not appear in Ohio City’s top-ten similar list.

Two things connect these. First, they cluster on exactly the categories the study flagged: transit, points of interest, schools, commute, walkability, character, and similarity. Second, none of them are hedged. The model states the invented station or the wrong commute with the same confidence it states a correct fact, which is precisely why a reader cannot tell the difference without the data behind it.

6. Why location is the input AI cannot fake

Models are closing the gap on almost everything. Location is different for three structural reasons, and the results above trace directly to them.

  • It is not in the training data. Schools, points of interest, and climate are where models hallucinate most, because the underlying facts were never reliably in the corpus. This is why the contradicted rate is highest on exactly these categories, and why grounding helps most there.
  • It cannot be crawled. The strongest live-search product collapses on “find a similar neighborhood,” because the answer needs a quantitative similarity computation over measured scores, not a page to read. This is why retrieval alone tops out below grounding.
  • It cannot be reasoned to. A model can infer that a place is “walkable” from context, but not the calibrated, comparable number, and it cannot tell you where its guess is wrong. For a ranking, a filter, or a real decision, the adjective is not enough.

The four walls of a home are on every website. Everything outside them, measured, calibrated, and comparable, is the part no model and no crawl reproduces on its own.

7. Limitations

Reported in brief; the full pre-registration and methodology notes are available on request.

  • The retired hallucination gate. One gate was retired before analysis because its assumed grounded rates were too optimistic. Hallucination is reported as a measured reduction, and the retired gate is disclosed rather than dropped.
  • The no-harm test is a proxy. It uses the entity tier of the main benchmark. The dedicated independent public-fact set was deferred from this run.
  • Deferred arms. A naive context-stuffing control is not part of this run. The live-web-search control was executed and is reported.
  • Bounded circularity. The study takes as given that Local Logic is authoritative for its own scores, which is table stakes for any data-licensing conversation, not something this evaluation set out to prove. What it measures is whether a model, handed authoritative data, uses it or keeps guessing. The hallucination measurement is independent: ungrounded error rates were re-checked against the open web with no Local Logic involvement, over several thousand claims per model.
  • Data-quality flags. The commute and points-of-interest categories carry known data-quality flags (endpoint ambiguity and a historical retrieval issue) and should be read with that caveat.
  • Single-shot. Runs used no retries or self-consistency, which is deployment-realistic but leaves some model variance uncontrolled.

8. What this means for teams building with AI

The practical conclusion is short. If you are putting AI in front of anyone making a location decision, the accuracy of the location layer is a product risk, and it is one you can remove. Grounding a model in verified location data is the lever, and the study shows it works across the full range of models, with the largest gains on the smallest and cheapest ones. That last point is the one to sit with: grounding is not only how you make the answer trustworthy, it is how you make a trustworthy answer affordable at scale.

The Local Logic MCP is how a model connects to that data. It exposes our verified location data as tools any MCP-capable model can call at the moment of the question, which is the exact configuration measured in this paper.


Appendix A: Benchmark coverage

Questions by category (490 total). affordability 90, commute 75, character 60, livability 45, demographics 45, climate 45, similar 24, comparative 24, transit 21, parks 21, points of interest 20, schools 20.

Scoring tiers. Deterministic: livability, climate, affordability, commute, demographics. Entity: transit, parks, schools, points of interest. Synthesis: character, comparative, similar.

Coverage. 46 neighborhoods across 39 cities, split roughly evenly across famous, mid, and obscure tiers by Wikipedia-pageview tercile, from high-traffic neighborhoods such as Williamsburg in Brooklyn and Silver Lake in Los Angeles to lesser-known ones such as Butchertown in Louisville and Bay View in Milwaukee. Roughly 20% are Canadian.

Scale. 7 paired models, 7,350 verified answers, 153,530 scored claims.

Appendix B: Generation cost and reliability

Cost per answer, from memory and grounded, with average tool calls in the grounded arm.

Model Cost, from memory Cost, grounded Avg tool calls (grounded)
GPT-5 $0.0289 $0.5354 4.7
Claude Opus $0.0581 $0.5174 4.0
Claude Sonnet $0.0059 $0.0736 2.8
DeepSeek Flash $0.0002 $0.0098 5.8
Gemini Flash $0.0007 $0.0063 1.8
GPT-5 nano $0.0014 $0.0039 3.7
GPT-4o mini $0.0002 $0.0040 1.8
Perplexity Sonar Pro (live search control) n/a $0.0160 n/a

Non-answer rates (refusals or retrieval-error stubs, reported separately and not graded) were low across arms, with one cheap grounded model the notable exception at about 8%.

Appendix C: Glossary

  • Grounding. Answering from verified external data supplied at the moment of the question, rather than from patterns in training. The opposite of answering from memory.
  • MCP (Model Context Protocol). An open standard for exposing data and tools to a language model so it can call them during an answer. The Local Logic MCP exposes our verified location data this way.
  • Verified-fact volume. The count of supported, checkable claims in an answer. A coverage measure, reported as a grounded-to-ungrounded ratio.
  • Claim verification rate. Supported claims divided by supported-plus-contradicted claims. The accuracy percentage used throughout.
  • Contradicted rate. The share of checkable claims that live data flatly contradicts. The primary hallucination measure.
  • Location-dependent categories. The categories where the answer depends on measured data: livability, affordability, climate, and demographics.
  • CI lower bound. The conservative end of a 95% confidence interval. Every headline number in this paper is reported at this bound.

This paper reports objective, measured attributes only. Demographic accuracy is treated strictly as a data-verification metric, not as a basis for characterizing who lives anywhere. All figures are from Local Logic’s July 2026 evaluation described above.

Frequently asked questions

What is grounding in AI?

Grounding means answering from verified external data supplied at the moment of the question, rather than from patterns in training. A grounded model looks a fact up instead of recalling it.

Does grounding reduce AI hallucinations, and by how much?

In this study, connecting models to Local Logic data cut the claim-level contradicted rate by about 2 to 4 times per model, and roughly halved answers with two or more serious errors, from about 32% to 17%.

Do cheaper AI models benefit more from grounding?

Yes. Accuracy lift was inversely related to the model standalone capability. The weakest model gained 16.6 points, and a grounded GPT-5 nano reached 97.6% on the product categories at roughly one-seventh the cost of full GPT-5.

Is web search a substitute for grounded location data?

No. A dedicated live web-search product scored 82% on the location-dependent categories, below every one of the seven grounded models. A calibrated, comparable location answer is not sitting on a page to be found.

What is the Local Logic MCP?

The Local Logic MCP is a hosted server that exposes verified location data as read-only tools any MCP-capable model can call during an answer, which is the exact configuration measured in this paper.

How was the study conducted?

Seven models each answered the same 490 neighborhood questions twice, once from memory and once grounded on the Local Logic MCP, with every claim fact-checked against live data. Every figure is reported at the lower bound of a 95% confidence interval.

Put grounding to work

The Local Logic MCP is how a model connects to this data in production. Explore the MCP solution page, browse the use case library, read the technical docs, or book a proof of concept.

Download Case Study