Accuracy
Accuracy answers one question: is AI telling the truth about you? Trakkr asks ChatGPT, Claude, and Gemini the questions your buyers ask, with live search switched off, then grades every answer against your own site, the facts you confirm, and an independently retrieved record of who you are.
Live search is off on purpose: what a model says from memory is what a buyer gets when they ask once and take the answer at face value, which is where a wrong answer sits unnoticed for months. Citations covers the search-first surfaces. Open Accuracy when a prospect quotes a price you do not charge, or when answers about you keep naming someone else.
The first run
A brand that has never been checked gets an activation screen. It names the three assistants, lists the buyer questions Trakkr drafted from what your company does (six by default, grouped as Discovery, Evaluation, and Diligence), and offers one button: Run the first check. The run happens in the background and survives a reload, and weekly monitoring starts after it. Viewers see the screen with no button.
The verdict band
The report opens with one badge, one sentence, and three readings.
| Badge | What it means |
|---|---|
| All clear | Nothing false, no open site disagreement, and full coverage |
| Worth a look | Nothing false, but a site disagreement or an unanswered question is open |
| Partial check | An assistant or an evidence file was missing, so this is not an all-clear |
| Needs attention | At least one answer contradicts a receipt |
The order is fixed: a real error outranks a site disagreement, which outranks incomplete coverage, so a run that could not read all of its evidence never wears the green badge.
The three readings are Wrong answers (with a small trend and the change since the previous check), Questions (measured, and answers graded), and Corrections (wrong answers that have since gone right, or your count of complete checks).
The question matrix
One row per buyer question, one column per assistant, sorted worst first. A cell is that assistant's verdict on that question at the last check.
| Verdict | The cell says |
|---|---|
| Checks out | Consistent with your site and the verified record |
| False | A material claim contradicts a receipt, and the error repeated on re-ask |
| Sends buyers elsewhere | Asked about you, the answer mainly recommends someone else |
| No answer | It says it does not know. Not wrong, but the buyer leaves empty-handed |
| Unverified | It makes claims nothing on file can settle yet |
Click a cell to open that assistant's answer, or the row to open the worst one. The Last 6 checks column carries the row's history, and the filter chips narrow the list to Wrong, Sent elsewhere, No answer, Unverified, or Checks out. With the list focused, j and k walk it, Enter opens a row, and r re-checks one.
The panel is yours. Add and check measures a new question straight away, up to 25 per brand, and Edit questions lets you reword one (which starts a fresh trail), switch it between weekly and monthly, or archive it. A question right eight checks in a row drops to monthly, and any false answer promotes it back to weekly.
Claims that nothing on file can settle collect in a row at the top of the band, each taking one tap: True, False (then what is true instead), or Doesn't matter.
Reading one question
Opening a row gives you the answer as the primary object, with tabs to switch assistants and the grading as annotation.
- The answer is quoted verbatim, exactly as that assistant returned it.
- What it claimed lists the material claims, each with its verdict and, for a red mark, its receipt.
- A receipt names its source: "You confirmed", "Your site says" with the quote and page, or "Verified record".
- An unsettled claim carries the Is this right? loop inline, and your answer saves to the verified record.
The footer offers Fix this on a false answer (it adds the finding to your action plan), Ask Agent, and Re-check now. The same list sits on the row's ... menu.
Where the truth comes from
The truth oracle settles five identity facts: the CEO, whether the company is public or private, the founders, the parent company, and the headquarters. Trakkr puts the pointed question to two independent live-search engines, normalizes each answer, and takes the majority cluster. No model decides what is true, clustering does. Both sources must agree before a value counts as high confidence, and Trakkr never accuses an assistant without one. The oracle is cached and refreshed about every two weeks.
That scope is deliberate: independent search suits identity facts and is weaker on plan names, prices, trial terms, integrations, certifications, and who a product is for. The site pass handles those. Your website is the authority for what you currently publish, so Trakkr reads the pages most likely to answer each buying question and compares the assistant claims with the page text. Prices, certifications, and company numbers are pulled out with deterministic rules and prices compared with arithmetic, keeping a model's memory out of the evidence.
The four site rows
| Row | What it means | What to do |
|---|---|---|
| Site differs | The assistants and your page state different values for one fact | Settle the live offer, then correct whichever side is stale |
| Site matches | The assistants repeat what your page publishes | Nothing. Keep it as evidence |
| Site silent | Buyers get an answer your page never gives | Publish the true answer where buyers expect it |
| Objection | A criticism your own copy contests | Treat it as buyer research, not an assistant error |
Site silent is the most useful row and the easiest to misread. It does not say the assistant is wrong; it says buyers are getting an answer your site does not support, and assistants are filling the silence from somewhere else.
Price disagreements show both sides on purpose. Calling $18 per user per month wrong against a page that says $20 per member per month would mean guessing the billing period, the region, and whether a user and a member are the same unit. Stating both sides stays true even when an annual discount later explains the gap.
The fix block
A proposed change appears as a diff: the line to remove, the line to add, and the page it belongs on. The removed line is never written by a model. It has to be an exact substring of an unexpired copy of the page saved for that check, and a near miss means the page changed underneath us, so nothing ships and the next crawl tries again. An edit is only offered when that line contradicts a fact you confirmed.
A Site silent row gets a draft instead: nothing is removed, it uses only confirmed facts, and nothing is published for you. A Site differs row never gets an edit, because that row declares no winner. Site findings that need an owner also appear as suggested work under the questions, ready for your action plan.
Corrections and check history
Corrections lists what stopped being wrong and when: the question, the assistant, the check that found it, the check that caught the change, and how many checks it has stayed right since. Only landed corrections appear; one that goes wrong again reopens.
Check history closes the report in two columns. On the left, every check newest first, each openable to see what moved since the one before (fixed, new, or regressed); a partial run is marked as not counted so it cannot flatter the trend. On the right, what this check did: how many prompts each assistant answered, the run's figures, and the way into three drawers.
- The verified record groups every fact by who settles it: verified independently, only your site can settle, or you confirmed.
- Buyer conversations holds the recorded multi-turn threads, replayed as the buyer saw them.
- The audit file is the whole check: how it ran, what each assistant said, which of your pages were read, and every claim with what happened to it.
Re-check, export, and print
Re-check reruns the whole check. It costs real provider money and takes several minutes, so manual runs are capped at one every seven days; while the cooldown is open the button says when the next one unlocks. Single questions are exempt: Re-check now in the drawer, or r on a row, remeasures just that one.
Export record opens the Accuracy Record, a print-tuned page at /accuracy/print with a native Print / Save as PDF. White-label is the default inside a client portal and an opt-in elsewhere. The menu beside it holds Download claims CSV and Copy summary. All three unlock once the report and its evidence are complete.
Limits Trakkr states out loud
- Every quote attributed to your site must be verbatim from a page Trakkr fetched. If it is not in the retrieved text, the row drops to Inconclusive.
- A dated page (a press release or news post) can show something happened. It can never establish a current CEO, price, owner, or offer.
- Site silent is only claimed for a question whose answering page Trakkr read. Anything outside the areas it covered stays unchecked, not missing.
- If your domain blocks every read attempt, the report says so and grades identity claims against independent sources only.
- A run that loses an assistant or an evidence file is saved as a partial check. Missing data is never read as correct.
Common questions
Why only three assistants when Trakkr tracks eight models?
Accuracy asks with live search off, and the point is what a model says from memory. AI Overviews, Perplexity, and the other search-first surfaces have no memory-only answer to grade, so they are covered elsewhere in the Visibility group.
Does a false verdict mean the assistant is definitely wrong?
It means a material claim contradicted a receipt (your page, a fact you confirmed, or the verified record) and the error reproduced when Trakkr re-asked. A contradiction that does not reproduce is never called false.
What happens when I answer "Is this right?"
Your answer joins the verified record and every future check grades against it, so the panel sharpens the more you correct it. You can undo it from the same place.
How often does the report update?
Checks land about a week apart on their own. Manual re-checks of the whole report are limited to one every seven days, while a single question can be re-checked at any time.
Why does one assistant show no answer for a question?
It did not answer on the last check. The drawer names that as a coverage gap rather than a clean bill, and the check-history band shows how many prompts each assistant answered.