Skip to main content
AI & Technology

How to Reduce AI Chatbot Mistakes About Prices, Policies and Products

Give your support chatbot clearer sources, safer lookup rules and a useful route to a person. Includes an answer-test matrix for prices, policies and products.

MattDarm10 min read
A customer-service adviser and developer checking product information beside a simple chatbot answer card
Illustrative chatbot review: verify the product source and the proposed answer before allowing the system to make a promise.

Key Takeaways

  • A chatbot should retrieve business facts from an approved source, not invent them from a plausible-sounding pattern.
  • Treat changing prices, stock and account details differently from general explanatory content. They may need a live, permission-controlled lookup.
  • Define what the bot must ask, refuse or pass to a person when information is missing, conflicting or outside its authority.
  • Test difficult questions and failure cases, not only the friendly examples used in a demonstration.
  • No setup guarantees perfect answers. Keep monitoring, human ownership and a way to pause the affected feature.

To reduce AI chatbot mistakes, start with the information and the decisions behind each answer. Identify which source is authoritative, how the correct record is selected and what should happen when the source does not provide a reliable answer.

A prompt saying “always be accurate” cannot fix an outdated price list, ambiguous product names or a tool that returns the wrong customer record. Nor can it give the bot permission to promise a discount or approve a refund.

This guide concerns a chatbot your business operates. Incorrect descriptions in public AI search are a different problem, covered in our article on AI getting your business wrong. Here, the business can control the support scope, connected sources, tests and escalation route. That makes a bounded improvement possible, even though perfection is not.

Start with the Answers That Could Cause Harm

Not all mistakes have the same consequences. A slightly awkward explanation of opening hours is different from promising a delivery date for a product that is unavailable. Prioritise questions where a wrong answer could cost money, mislead a customer or expose information.

Make a short list with the support team. Include prices and tax treatment, product compatibility, availability, delivery promises, cancellation conditions, account-specific information and any claim that requires professional judgement.

For each topic, write the permitted behaviour. The bot might explain a published policy, retrieve a price for a specific variant or collect details for a human review. Those are different powers. Do not give it a broad instruction to “resolve the issue” without defining what resolution may involve.

A useful AI chatbot development brief makes these boundaries visible before the interface is designed. A narrow assistant that handles routine questions accurately is often more useful than a confident assistant that tries to answer everything.

Create a Source Register

List the information the chatbot is allowed to use. For every source, record its owner, purpose, update route and limitations. Do not assume that a folder of documents forms a coherent knowledge base simply because all the files were uploaded successfully.

Question typeSuitable starting sourceImportant restriction
General service explanationApproved public service contentDo not add services not offered
Product priceCurrent catalogue or pricing serviceMatch the correct product, variant and market
Delivery estimateApproved fulfilment informationDistinguish an estimate from a guarantee
Returns policyCurrent published policyDo not invent exceptions or approve a refund
Customer order statusAuthorised account lookupVerify access and minimise disclosed information

This is a planning framework, not a claim that every platform supports these connections without development. Check the actual tool, permissions and data model.

Microsoft's Copilot Studio capability guidance describes generative answers based on connected knowledge and notes the importance of source quality. The wider lesson is practical: the answer can only be as useful as the relevant information the system can retrieve and interpret.

Keep Exact Facts out of Guesswork

Prices deserve particular care. A product name may refer to several sizes, subscription options or trade arrangements. A PDF may contain an old list while the website displays a current offer. If the bot chooses the wrong record, better wording will not rescue the answer.

Ask for the missing identifier when it matters. That might be a product code, size, quantity or delivery location. Use the approved system to obtain the result, and validate that the returned record matches the request before displaying it.

Where a calculation is required, use tested business logic rather than asking a language model to improvise it. Define how tax, shipping, minimum quantities and discounts are handled. If the available information is insufficient, offer the correct quotation route instead of producing a confident estimate dressed up as a final price.

Keep a boundary between explaining a price and committing the business. An assistant may show an approved catalogue price while a bespoke quote still requires a person to check scope. Make that distinction clear to the customer.

Remove Conflicting and Unnecessary Sources

More documents do not automatically mean better answers. A chatbot can struggle when several sources describe the same subject differently and nobody has declared which one takes precedence.

Remove obsolete material from the active retrieval set where appropriate, or label it clearly if it must remain available for a specific purpose. Separate internal discussion notes from approved customer-facing policies. A draft pricing conversation should not become an answer to a visitor's question.

Write a source-precedence rule in plain language. For example, a current product record may take priority over an archived brochure. Then test whether the implementation follows that rule when both are retrieved. Do not rely on the rule being buried in a long general prompt.

Changing the source library is itself a controlled change. Keep enough version information to understand which source was available when a disputed answer was produced. Our broader AI-powered customer service work includes the human operating process around the bot, not just the conversation window.

Configure Real Tools, Not Imaginary Capabilities

An instruction to “check the order system” is meaningless if no order-system connection exists. It can also encourage a misleading answer that sounds as if a check has been performed.

Microsoft's guidance on grounding instructions in configured knowledge and tools makes this distinction explicit. Describe only the actions and sources the implementation actually supports.

Test the unavailable case. If a lookup times out, returns no record or rejects access, the bot should explain the limitation without exposing technical details or fabricating a result. It should not turn “could not check” into “not in stock” or “your refund has been approved”.

Make successful tool use observable to the support team through appropriate logs. A polite answer is not proof that the tool was called. Equally, a successful tool response is not proof that the answer used the correct fields. Check both stages.

Write a Useful Uncertainty Response

A fallback should help the customer continue. “I cannot help with that” may be safe but frustrating if it provides no route forward.

For a missing product identifier, ask a focused question. For conflicting policy information, explain that the team needs to confirm the answer. For an account issue, direct the customer through the approved authentication or support route.

Do not ask for unnecessary sensitive information in an open chat. Collect only what the next step genuinely requires and through the appropriate channel. A handover summary can describe the issue without copying every personal detail from the conversation.

For example, an illustrative fallback could say: “I cannot confirm the price for that configuration from the current catalogue. I can help you send the product code and required quantity to the team for a checked quote.” This is useful without pretending that the business has made a commitment.

Build an Answer-Test Matrix

Choose representative questions and write the expected behaviour before running the bot. Include ordinary requests, ambiguous requests and situations where the correct result is to stop.

Test caseWhat a good response doesWhat should fail the test
Product name matches two variantsAsks which variant is intendedPicks one without clarification
Old brochure conflicts with current catalogueUses the approved current source or escalatesCombines both prices
Customer asks for an unapproved discountExplains the agreed route for approvalPromises a discount
Lookup service is unavailableStates the limit and offers a next stepClaims the lookup succeeded
User asks it to ignore the policyKeeps the business boundaryTreats the user's instruction as authority

Use fictional or appropriately sanitised test data. Run the tests after changes to prompts, tools, sources or models. Include follow-up questions, because a first answer can be correct while a later response drifts away from the source.

Do not let the bot grade itself as the only quality check. Human review should examine whether the answer is supported, relevant and within authority. Keep disagreements visible so the test set improves rather than silently passing doubtful cases.

Measure Reliability, Not Only Deflection

A high number of conversations ending without a human is not automatically a success. Some customers leave because they received an answer; others leave because they gave up.

Review a sample of conversations alongside escalation outcomes and complaints. Look for unsupported promises, missed clarifying questions, repeated dead ends and cases where a human had to correct the answer later.

Microsoft's monitoring FAQ explains that automated answer-quality analysis uses samples and has limitations. Treat such tools as aids to review rather than unquestionable verdicts.

Keep your own measures understandable: unsupported-answer rate in the reviewed sample, correct escalation behaviour, source retrieval failures and the time needed to resolve a handed-over enquiry. Define the sample and avoid presenting it as a perfect measure of every conversation.

Assign an Owner and a Stop Button

Someone needs authority to pause an unreliable topic or connection. That person should know how to remove an obsolete source, restore a previous configuration and notify the team about affected answers.

The operating plan should also say what happens outside working hours. A bot should not promise an immediate human response when nobody is available. Set expectations that match the real support arrangement.

Our article on the customer backlash against poor chatbots explains why the quality of the handoff matters. Customers should not have to repeat a long conversation because the automation reached its limit.

If you are reviewing an existing assistant, start with ten questions that could create the most expensive misunderstandings. Gather the approved answers, test the current bot and record the gap. An AI strategy review can then prioritise the changes by risk and usefulness. Fix the source and the decision boundary before spending another week polishing the greeting.

Frequently Asked Questions

Can we stop an AI chatbot making mistakes completely?

No setup guarantees perfect answers. You can reduce risk with approved sources, validated lookups, narrow permissions, test cases, monitoring and human escalation. Keep a way to pause unreliable behaviour and investigate the underlying cause.

Should a chatbot calculate prices itself?

Use approved pricing data and tested calculation rules for exact commercial figures. The bot should clarify missing product or order details and avoid presenting an unsupported estimate as a final quotation.

Will uploading more documents make the answers better?

Not necessarily. Duplicate, old or conflicting documents can make retrieval harder. Organise the source set, identify the authoritative version and test whether the system uses it correctly.

What should the bot do when a tool or data source fails?

It should acknowledge the limitation and offer an appropriate next step, such as clarification or a human handover. It must not claim a lookup succeeded or invent a business result from a technical failure.

How should we test a customer-service chatbot?

Prepare expected outcomes for normal, ambiguous and failure cases. Check source support, product matching, permissions and escalation across follow-up questions. Re-run the tests after changes and include human review rather than relying only on automated scores.

AI ChatbotsCustomer ServiceAI ReliabilityKnowledge Bases

Share this article

Stay ahead of the curve

Weekly insights on web development, AI, branding & digital marketing. No spam, unsubscribe anytime.

By subscribing you agree to our Privacy Policy. Unsubscribe at any time.

Adam Saez
Alina Stefanovičiūtė
Daniel Ashby
Matt Laybourn
Richard Jones
Paul Campbell

People we've worked with

Real projects, built together

Let’s Grow Your Business Together

Tell us about your project and we’ll show you exactly how we’d grow your business. Book a free 30-minute discovery call, no pressure.