How Lodestar scores.
Every published score has to be re-derivable from the method below and the raw artifacts we publish with it. Change the method and the version changes; old scores keep their stamp.
Pilot edition Chicago HVAC · Aug 2026 — scores from a multimodel visibility battery and a readiness probe. How scoring works
The short version
We ask AI assistants the questions real customers ask, in a battery we publish. We record who gets named. Separately we probe each business's public surface — twice — to see whether machines can read it at all. Those two numbers, plus ground truth once it is fielded, make the score.
Reliability is a published number, not a promise: this edition's probe flip rate is 1.7%, against a hard publish gate of 5%. An edition that fails the gate is held and re-probed, never shipped.
Edition notes — draft-v0.7-multimodel-pilot
These are the caveats that travel with this edition. They are deliberately here rather than on the front page, but nothing is hidden: the same text ships in /index.json and /llms.txt.
DRAFT — not a final public Lodestar Score. Visibility = multimodel pilot: ChatGPT 20 + Gemini 20 + Perplexity 20 = 60 cells; Claude not fielded. Gemini surface caveat: PILOT-01–03 logged-in Work vs 04–20 unsigned Flash-Lite. Perplexity unsigned Free. Readiness = stabilized dual-pass probe (v1.0.1). Ground truth pending.
- Edition id:
draft-v0.7-multimodel-pilot - Visibility (40): multimodel pilot battery (ChatGPT+Gemini+Perplexity, 60 cells)
- Readiness (35): 0-25 of 35 pillar from probe (robots, llms.txt, sitemap, schema.org, MCP verify)
- Ground truth (25): 0 of 25 — pending
- Fielded: 2026-08-12 multimodel: ChatGPT 20 + Gemini 20 + Perplexity 20 = 60 cells; Claude not fielded. Gemini: PILOT-01–03 logged-in Work vs 04–20 unsigned Flash-Lite. Perplexity unsigned Free.
- Probe: 2026-08-11 (dual-pass stabilized)
- Models fielded: ChatGPT, Gemini and Perplexity Claude not fielded this edition.
- Generated: 2026-08-12T03:16:29.142389+00:00
Everything on this page in machine-readable form: /index.json and /llms.txt, plus the Trust API at api.lodestarindex.com.
Lodestar Index — Methodology v1.0 (draft for comment)
Status: Draft v1.0 · August 10, 2026 · This document becomes the public /methodology page on lodestarindex.com. Every published Lodestar score must be re-derivable from what is written here. When the method changes, the version number changes, and old scores keep their version stamp.
1. What we measure
A Lodestar score (0–100) answers one question: how findable, readable, and trustworthy is this business to AI agents acting for a customer? It is built from three pillars:
Visibility (40 points). When AI assistants are asked the questions real customers ask, does this business get recommended — and how often, and how prominently?
Readiness (35 points). Can automated agents actually read and interact with the business's digital surface? Crawl permissions, machine-readable summaries, structured data, and (increasingly) transactable agent endpoints.
Ground truth (25 points). Is what an agent would learn accurate? License status verified against official registries, review-pattern integrity, and consistency of the business's core facts across the web.
v1.0 measures Visibility and Readiness fully and Ground Truth partially (license verification + review-pattern heuristics + name/address/phone consistency). Deeper ground truth (pricing honesty, booking follow-through) arrives in v2 and will be version-stamped when it does.
2. The query battery
Each vertical/metro pair gets a battery of ~100 queries across five intent categories, drawn from how real customers actually ask:
- Discovery — "Who's the best HVAC company in Evanston?"
- Problem-driven — "My furnace is blowing cold air — who should I call near Oak Park?"
- Comparison & reputation — "Is [company] reputable? Who's better, X or Y?"
- Transactional — "Book me a furnace tune-up near 60614." (This category is weighted highest: it's where money changes hands.)
- Constraint-based — "Emergency 24/7 furnace repair with financing in Naperville."
Queries are parameterized by neighborhood/suburb and rotated. ~70% of the battery is published; ~30% is held out and rotated each cycle so scores cannot be gamed by optimizing against a fixed list.
Models tested: ChatGPT, Claude, Gemini, and Perplexity — the consumer products, because that is what customers actually use. Each query runs 3 times per model per cycle (AI answers vary; we measure recommendation frequency, not a single lucky answer). We capture: every business mentioned, its position, qualifying language ("highly rated," "mixed reviews"), and the sources the model cited.
3. Scoring rubric v1.0
| Pillar | Component | Points |
|---|---|---|
| Visibility (40) | Recommendation share across the battery, weighted by intent category (transactional highest) | 30 |
| Prominence & sentiment of mentions | 10 | |
| Readiness (35) | Agent/transactable surface (MCP endpoint, machine-readable booking) | 15 |
| AI crawl policy (explicit robots.txt treatment of AI crawlers) | 5 | |
| llms.txt / machine-readable site summary | 5 | |
| Structured data (schema.org business markup) | 5 | |
| Sitemap & basic discoverability | 5 | |
| Ground truth (25) | License verified against official state/municipal registry | 10 |
| Review-pattern integrity (velocity anomalies, distribution shape) | 10 | |
| NAP consistency (name/address/phone identical across major surfaces) | 5 |
Scores publish as 0–100 with a letter grade. Component scores are visible on every business's page — a business should never wonder why it scored what it scored.
4. Reproducibility standard
For every published cycle we publish: the query list (public portion), capture dates, model names and versions, run counts, and per-business raw mention counts. Anyone with the same tools can re-run our public battery and get materially the same results. Claims we cannot make reproducible, we do not publish.
The publish gate: a trust index that flips on re-probe is not a trust index. Every signal is probed twice with agreement required before it publishes, and no market edition ships while its measured flip rate is 5% or higher. The flip rate itself is published with every cycle — our reliability is a number you can check, not a promise you have to take.
5. Cadence & versioning
Full query battery: monthly per market. Readiness probes: weekly. Ground-truth checks: monthly. Scores carry their date and methodology version forever. The methodology itself is versioned semantically: component weight changes bump the minor version; pillar changes bump the major version; every change is logged publicly.
6. Anti-gaming
Held-out rotating queries (30%), outcome-weighted scoring (being genuinely recommended matters more than checkbox compliance), anomaly detection on sudden score jumps, and the founding rule that makes gaming pointless at the source: payment never touches the grade. No paid placement, no lead-gen, no exceptions.
7. Fairness & corrections
Any business can claim its listing free. Disputes with evidence get investigated within one cycle. A business that fixes a deficiency gets re-tested free in the next cycle — improvement should show up fast. Corrections and their reasons are logged publicly.
8. Known limitations (v1.0)
Consumer AI answers vary by session, location, and account history — our 3-run sampling approximates, not exhausts, this variance. Readiness signals are proxies for agent usability, not guarantees. Ground truth v1 covers license status, review-pattern heuristics, and NAP consistency only — it does not yet verify pricing or service quality. We publish these limitations because a trust index that hides its own uncertainty doesn't deserve the name.
Lodestar Index · lodestarindex.com · Methodology v1.0 draft · Comments to preston@banjo.ventures
Methodology Addendum v1.0.1 — Probe Reliability Protocol
Aug 10, 2026 · Adopted after an independent external stress test re-ran our day-one probes and observed ~28% signal instability. A trust index that flips on re-probe is not a trust index. This addendum defines the stabilization protocol; it merges into methodology v1.1.
Why probes flip (observed causes)
Soft-404s (sites returning a 200 homepage for any path, making llms.txt look "present"); WordPress plugins serving llms.txt intermittently; sitemaps that exist at standard paths but aren't declared in robots.txt; WAFs and timeouts blocking probes; redirect chains; classifier edge cases on non-standard directives. Some instability is genuinely site-side variance — which is itself a finding — but we only publish what is stable.
The stabilization protocol (all published signals)
- Two-probe agreement. Every signal is probed at least twice, hours apart. Published value = the agreeing value. Disagreement → a third probe; still unstable → published as "variable" (its own state, scored as not-present but labeled honestly).
- Soft-404 detection. A file "exists" only if its content is plausibly that file type: llms.txt must be text/markdown-like and must materially differ from the homepage; a response matching the site's known-404 or homepage fingerprint = absent.
- Sitemap multi-path. Checked in robots.txt declaration AND at standard paths (/sitemap.xml, /sitemap_index.xml). Existing-but-undeclared = present, sub-noted.
- MCP verification, not MCP advertisement. An agent endpoint counts only if it returns valid, schema-conformant JSON at a standard path (or a working declared endpoint). Text that merely mentions MCP tools = "advertised — unverified," never "offers."
- Unreachable ≠ zero. Fetch failures publish as "unreachable," excluded from score denominators, retried next cycle.
- Retry discipline. Transient failures get one retry after a delay before classification.
- Published classifier. Exact classification rules and thresholds for every signal state publish with methodology v1.1 — a business must be able to derive its own classification.
The publish gate (non-negotiable)
No market edition publishes while its measured signal flip rate is 5% or higher across consecutive same-week probe passes. This is a gate, not a target: an edition that fails it is held, re-probed, and fixed — never shipped. The flip rate itself is measured every cycle and published alongside the scores — our reliability is a number, not a promise.
A trust index that flips on re-probe is not a trust index.
Also adopted from the stress test
- Readiness pillar completion: structured-data (schema.org) probing and MCP endpoint verification join the probe set (the two of five readiness signals not yet fielded).
- The current Aug 10 dataset is marked DRAFT in every artifact pending a stabilized re-probe.
- All published third-party statistics carry source attribution inline.
- Per-company scorecard pages ship before the 58 scorecard emails go out.
- Canonical pricing lives in the financial model (monitoring $149/mo · agency $500–1,500/mo · certification $1,500/yr); other documents sync to it at next revision.
“Payment never touches the grade. A trust index that flips on re-probe is not a trust index. The score can't be bought — only earned.”
Found something wrong? Send evidence to preston@banjo.ventures. Corrections are re-checked in the next edition and the method is versioned in public.
Where to next
The scores this method produced, for every business in the market.
Open any scorecard and see the signals behind the number.
Why payment never touches the grade.