Your AI chat dashboard says 3 unanswered questions out of 40 conversations. That is not 3 wrong answers out of 40. It is 3 conversations where the assistant admitted it did not know something.

The assistant that told a visitor your Pro plan includes SSO, when it does not, is not in that number. It answered confidently, so nothing flagged it. That is the failure that loses the customer, and the one metric you are given cannot see it.

The unanswered questions counter flags a conversation when an assistant message contains one of five hardcoded refusal phrases, so it measures admitted ignorance, not accuracy. Confident wrong answers score zero. To catch those you read transcripts: filter to AI-only conversations, read the visitor's question before the answer, and check every claim about price, plan, or availability against your actual page.

What the unanswered questions counter actually counts

OperatorStack computes it by scanning assistant messages for five literal strings:

i'm not sure
i don't have information
i can't help with that
you might want to contact
i don't know

The match is a case-insensitive substring check against the message body. Three details change how you should read the number:

  1. It counts conversations, not messages. A conversation that trips the check eight times counts once. Ten bad conversations and one catastrophic one look the same.
  2. It only looks at assistant messages. A visitor typing "I don't know" never affects it.
  3. The window starts at your date range but the check is per conversation. A conversation created inside the range is scanned across all of its assistant messages.

So the counter answers one narrow question: how many conversations did the assistant explicitly bail out of? That is worth knowing. It is not a quality score.

If you have written a custom system prompt, this counter is probably reading zero for the wrong reason. The five phrases match OperatorStack's default refusal wording. Tell your assistant to say "that is outside what I can help with" and the substring check finds nothing, so the dashboard reports zero unanswered questions while your assistant refuses visitors all day. Nothing warns you that the metric went blind.

The three failures nothing flags

Reading transcripts is how you find these. There is no counter for any of them.

Confident and wrong. The assistant states a price, a plan limit, or an integration that does not exist. This is the expensive one. The knowledge base loads your page text into the system prompt, so an outdated pricing page produces a fluent, specific, wrong answer.

Right answer, wrong question. The visitor asks "can I export to Postgres" and the assistant answers about CSV export. It is not wrong about CSV. It did not answer the question, and the visitor leaves thinking the answer was no.

Stale. You changed your pricing three weeks ago and never rescanned. The assistant is accurately quoting a page that no longer exists.

A 20-minute weekly review

1

Filter to AI-only conversations. On the AI Chat page, open the Chats tab. Conversations carry a mode of ai, live, or handoff_to_ai. Only ai is the assistant working alone. If you judge the bot on a conversation you took over halfway, you are grading your own typing.

2

Read the visitor's message before the answer. Read them in that order, deliberately. Reading the answer first makes it sound reasonable, and you will miss that it addressed a different question. Every message has a role of user or assistant, sorted by created_at.

3

Check every factual claim against the real page. Price, plan limits, availability dates, integrations, refund terms. Open your pricing page next to the transcript. Anything the assistant asserted that you cannot find on your own site is either invented or stale, and both need the same fix.

4

Tag each bad answer with one of three causes. Missing knowledge (the answer is not on your site anywhere), stale knowledge (it is on your site but the assistant has an old copy), or out of scope (the assistant should have declined). The cause decides the fix, and the three fixes are completely different.

Pulling transcripts instead of clicking through them

The Chats tab is fine for ten conversations. At a hundred it is slow, and you cannot search what visitors actually typed. Open your browser console on the dashboard and call the same endpoints it calls:

// List AI-only conversations, newest first.
const res = await fetch("/api/v1/chat/conversations?per_page=100&mode=ai", {
  credentials: "include",
});
const { items, total } = await res.json();
console.log(`${items.length} of ${total} conversations`);

Then pull one transcript and print it:

const detail = await (
  await fetch(`/api/v1/chat/conversations/${items[0].id}`, {
    credentials: "include",
  })
).json();

detail.messages.forEach((m) => console.log(`[${m.role}] ${m.content}`));

The detail response carries messages sorted oldest first, plus tool_executions if the assistant saved a contact or signed someone up mid-conversation. To grep every transcript for a word visitors actually typed:

const hits = [];
for (const c of items) {
  const d = await (
    await fetch(`/api/v1/chat/conversations/${c.id}`, { credentials: "include" })
  ).json();
  if (d.messages.some((m) => m.content.toLowerCase().includes("refund"))) {
    hits.push({ id: c.id, summary: c.summary });
  }
}
console.table(hits);

You need that loop because the search parameter on the conversations list does not do what it looks like it does. It runs an ilike against the conversation summary field, which is AI-generated, not against message content. Searching for refund returns conversations whose summary happens to mention refunds and silently misses every conversation where a visitor typed the word but the summarizer wrote something else.

Fixing what you found

The cause you tagged in step 4 maps to exactly one fix:

Missing knowledge goes in as an FAQ entry. FAQ entries load first in the knowledge base priority order, ahead of scanned page text, so an FAQ answer wins over whatever the crawler picked up. This is also the fastest fix: one entry, no rescan.

Stale knowledge needs a rescan. The crawler holds a copy of your pages from whenever it last ran, and nothing re-crawls automatically. Changed your pricing page? The assistant does not know until you trigger a new scan.

Out of scope is a system prompt problem, not a knowledge problem. Adding more knowledge will not stop an assistant from answering questions it should decline. That is what guardrails are for, and it is the one fix where adding content makes the problem worse.

Fix the stale ones first even though they feel least urgent. Missing knowledge produces a refusal, which is visible and merely unhelpful. Stale knowledge produces a confident, specific, wrong answer, which is invisible and actively misleads the visitor.

What to check monthly instead of weekly

The topics list on the AI Chat analytics is built from conversation summaries, not raw messages, and only conversations that have a summary contribute. It is a decent read on what visitors keep asking about and a bad read on any single conversation. Use it monthly to spot a question you should answer on the landing page itself, rather than letting the assistant field it forty times.

If one topic keeps appearing, that is not a chat problem. That is your homepage failing to answer an obvious question, and the assistant is absorbing the cost.

Frequently Asked Questions

How do I know if my AI chat is giving wrong answers?

You read the transcripts. No metric in the dashboard detects a wrong answer, because detecting one requires knowing the right answer. The unanswered questions counter only catches the assistant admitting it does not know, which is the failure mode that costs you the least.

What does the unanswered questions number actually count?

It counts conversations, not messages, where at least one assistant reply contains one of five hardcoded phrases: i'm not sure, i don't have information, i can't help with that, you might want to contact, or i don't know. The match is a case-insensitive substring check, and a conversation is counted once no matter how many times it trips.

Can I search chat transcripts for a specific word?

Not the message text. The search filter on the conversations list matches the AI-generated summary field only, so searching for "refund" finds conversations whose summary mentions refunds, not every conversation where a visitor typed the word. To search message content you have to pull the transcripts and search them yourself, as in the console snippet above.

Does writing a custom system prompt change the unanswered count?

Yes, and silently. The five phrases are matched against OperatorStack's default refusal wording. If your custom prompt tells the assistant to say "that is outside what I can help with" instead, none of the five substrings appear and the counter reads zero while the refusals keep happening.

How often should I review AI chat transcripts?

Weekly while you are still changing your pricing, positioning, or product copy, since those are what the assistant gets wrong after a site change. Twenty minutes covers a normal pre-launch week. Once the content settles, monthly is enough, plus a spot check after every site update.