Skip to content
Planning a Business AI Chatbot: Scope, Data & Evaluation

Planning a Business AI Chatbot: Scope, Data & Evaluation

A business AI chatbot is a product decision, not just a prompt. It can answer questions, help a person find information, guide a workflow or prepare an action. To be useful, it needs a defined audience, an approved information boundary, a clear way to handle uncertainty and a process for checking quality after launch. Without those pieces, a fluent answer can still be wrong, incomplete, unsafe or irrelevant to the task the business intended to improve.

This guide helps owners and project teams decide whether a chatbot fits a real business need and how to plan one responsibly. It walks through use-case selection, source material, privacy, security, answer design, evaluation, human review, integration and ongoing operations. It is a planning guide, not legal advice, a security certification or a guarantee that a particular model will behave correctly. Requirements differ with the data, users, jurisdiction and actions involved. Use qualified privacy, legal and security advice where the risk calls for it.

1. Begin with a recurring task, not a model name

Start by describing the person who needs help and the task they are trying to finish. “We want an AI assistant” is too broad to evaluate. A more useful statement might be: “A customer wants to find the current delivery instructions in our approved help material, understand the next step and reach a person if the answer does not fit.” The statement names a user, a question, a source boundary, a result and a fallback. It can be tested before the team chooses a vendor or interface.

Look for evidence that the task occurs often enough to matter. Review support themes, existing documentation, search queries on the site, call notes or staff observations, while protecting personal information. Separate repeatable questions from cases that need judgement, empathy or access to a person’s unique record. A chatbot can help someone locate an answer that already exists; it cannot make an unclear policy coherent or substitute for an owner who has not decided what the policy is.

Estimate the current effort and the desired change. Are staff copying the same approved instructions into email? Do visitors struggle to find a particular page? Is a process delayed because required information is scattered? Describe the baseline qualitatively or with a measured internal figure if one is available. Do not invent savings or conversion improvements before a pilot. Decide what evidence would make the project worth continuing and which harms would make it unacceptable, such as exposing another user’s details or confidently giving an unsupported instruction.

Chatbot use-case flow from a recurring visitor question through approved information, uncertainty and a safe response
Original Heavenkeys use-case diagram. It illustrates a planning pattern, not measured chatbot performance. NIST’s Generative AI Risk Management Framework profile provides a broader risk-planning reference.

2. Choose a use case where a conversational interface helps

Conversation is useful when people ask the same underlying question in different words, need help navigating a collection of information or benefit from a short sequence of clarifying questions. It is less suitable when the task is already faster with a clear form, a table or a direct link. A chatbot should earn its place in the experience. If it adds an extra step, hides the contact route or forces a visitor to ask for information that was already visible, it has made the service harder to use.

Compare the proposed assistant with a simpler alternative. Could an improved search page, clearer navigation, a concise FAQ, a guided form or a better account dashboard solve the problem with less uncertainty? A conversational interface may still be valuable, but the project should explain the reason. The team should be able to state what the assistant does that the alternative cannot do as well, what information it needs and how the user can leave the conversation.

Favor tasks with bounded inputs and a verifiable finish. Finding a published policy or explaining the fields in a public form is easier to evaluate than making an open-ended business decision. Avoid starting with high-impact advice or irreversible actions. If the assistant can affect eligibility, finances, health, employment, legal decisions or access to an essential service, the team needs a careful risk and governance assessment before deployment. A general-purpose chatbot plan should not be used to bypass those obligations.

3. Draw the boundary around the assistant’s job

Write a short scope statement for what the assistant may do, what it must not do and what it should do when a request falls outside the scope. Examples of allowed work might include summarizing an approved public article, locating a service page or explaining a published process. Restricted work might include making a promise on behalf of staff, changing a customer record without permission, presenting a guess as policy or making an individualized decision that requires a qualified person.

Define the audience and the context. A public website assistant may answer from public information, while an internal employee tool might use controlled company material. A logged-in customer portal can involve personal records and authorization checks. Those systems have different privacy and security boundaries. Being signed in is not sufficient proof that a user may access every record or take every action. The application must enforce the relevant permission at the data and action layer.

Make the boundary visible in the user experience. Explain what the assistant can help with and when a person should contact the organization directly. Do not make visitors infer whether the answer is official or generated. Use an explicit disclosure appropriate to the interface and make escalation easy to find. The correct wording depends on how the service works; avoid claiming a human reviewed an answer unless that actually happened.

4. Prepare approved source material before connecting a knowledge base

Retrieval systems can help an assistant find information from a selected set of documents, but retrieval does not make the documents accurate. Start by identifying the source owner, approval status, effective date and intended audience of each file. Remove duplicates, obsolete policies and draft material that should not appear to users. Establish how changes are reviewed and how an outdated document is removed from search. If two sources disagree, the assistant needs a clear precedence rule or should disclose that the information is unresolved and route the question to a person.

Use source material that is specific enough to answer the intended task. A folder containing years of presentations, meeting notes and policy drafts may look comprehensive while mixing different versions and audiences. Prepare a smaller, governed collection first. Include the document title and an appropriate link with retrieved passages so the answer can point the user to its basis. A link should resolve to information the intended user is allowed to see; the assistant should not expose a restricted internal source through a public response.

Identify missing information and unclear terms as content problems, not prompt problems. If the source never defines who approves a request, adding instructions to “be helpful” will not establish the policy. Ask the accountable business owner to resolve the ambiguity and approve the wording. Maintain a change log for source updates, including who approved them and which evaluations should be rerun. This connects content governance to system behaviour and gives the team a way to understand why an answer changed.

Knowledge flow from approved current sources through retrieval and grounded answers to source review and gap tracking
Original Heavenkeys knowledge-flow diagram. The source collection needs its own owner and update process; use NIST’s AI Risk Management Framework resources to structure broader risk conversations.

5. Protect source data and user information

Map the information the assistant receives, retrieves, stores, sends to a model provider, includes in logs and passes to connected services. Include the conversation text, user identifiers, retrieved documents, telemetry, support transcripts and any action payload. Ask who can access each data store, how long information is retained, how it is deleted and which providers or subprocessors receive it. These questions should be answered for the actual architecture, not inferred from a marketing description of a product.

Minimize the information collected. If a public help assistant can answer without asking for a full name, account number or sensitive detail, do not request those details. If the task requires account context, authenticate the user through the appropriate application and retrieve only the record needed for the current task. Avoid copying more customer data into a prompt than the model needs. Redaction can help in some situations, but the team must understand what is removed, whether identifiers remain and how the transformation affects the workflow.

Document the privacy and retention decisions with the responsible owner. Explain to users what information is processed and why, using language suited to the product and applicable requirements. Establish how a person can request support or exercise relevant rights. A chatbot transcript can contain new sensitive information even if the source documents are public. Do not assume that disabling a history screen deletes provider logs or downstream records; verify data handling against the contracts and technical controls in use.

6. Enforce access control before the model sees retrieved content

When a chatbot uses private information, permissions must be checked before restricted content enters the model context. It is not enough to tell the model “do not reveal private data” after retrieving it. The application should determine what the user may access, filter the available records and validate any action at the server or service layer. The model’s answer is not an authorization mechanism.

Test direct requests for another customer’s records, guessed identifiers, shared accounts, changed roles, expired sessions and access through a copied link. Include cases where a user asks the assistant to repeat or summarize information from a previous conversation. Clarify whether context persists between sessions and how it is separated across users. If a staff member has elevated permissions, verify that the interface and logs make the identity and authorized action clear.

Use the least privilege needed for integrations. A read-only source search should not carry a credential that can delete records. If the assistant can take an action, require a deterministic application endpoint with explicit validation, narrow scopes and an audit trail. The model can help draft or select a proposed action, but the application should decide whether it is allowed and whether user confirmation is needed. OWASP’s prompt-injection prevention guidance is a useful starting point for identifying classes of security risks, but it does not replace a threat model for your system.

Access-control diagram checking identity, permissions, filtered retrieval and action authority before an AI assistant can respond
Original Heavenkeys security diagram. OWASP’s prompt-injection prevention guidance covers security controls; apply them at the application layer for your own system.

7. Design how the assistant responds when evidence is missing

Decide what counts as enough evidence to answer. The rule can depend on the task: a public navigation question may need a relevant page link, while a policy answer may need a current, approved source passage. If retrieval returns nothing, conflicting material or an outdated page, the assistant should not fill the gap with a confident guess. It can ask a specific clarifying question, explain that it cannot verify the detail or route the conversation to the right person.

Make uncertainty useful instead of vague. “I do not have the current delivery window in the information available here” is more actionable than an invented estimate. Give the user a next step, such as a link to the official page or a contact route. If the request needs a person, explain what information to include and whether the conversation will be passed along. Do not imply that a handoff occurred until a system confirms that it was received.

Write rules for unsupported claims, safety-sensitive subjects and requests beyond the assistant’s role. The assistant should distinguish public general information from individualized advice. Where the user’s question could involve high stakes, a suitable response may be to stop and direct them to a qualified human or authoritative service. The correct boundary is determined by the organization’s service and obligations. Test the boundary with realistic examples instead of relying only on a short instruction in the system prompt.

8. Decide which tasks need confirmation or human review

Separate answering from acting. Summarizing a public page is different from sending a message, booking an appointment, modifying an account or issuing a refund. For each action, document the user’s authority, the data the system must validate, the confirmation screen and the recovery option if the result is wrong. Prefer a preview that states exactly what will happen before an irreversible or externally visible action occurs.

Human review can be an escalation path, an approval step or a quality sample; those patterns are not interchangeable. A review queue should specify who responds, what evidence they see and how quickly the user can expect a reply. If the assistant prepares a draft for a person, label it as a draft and show the underlying information. If a human is expected to approve every output, ensure staffing and service expectations make that feasible rather than promising review that cannot be delivered.

Define situations that require a handoff before implementation. These might include a missing policy source, repeated misunderstanding, account-specific data, a complaint, a safety issue or a requested transaction that needs authorization. Make the handoff work in the actual channel and during the hours the organization supports. When a person is unavailable, provide an honest alternative, such as a form or expected callback window, instead of leaving the visitor in a conversation that cannot be resolved.

Safe uncertainty route with four outcomes: answer from evidence, ask a clarifying question, explain what is unknown, or transfer to a person
Original Heavenkeys fallback diagram. It is a product-design pattern, not a claim about model accuracy. NIST’s GenAI risk profile discusses risks to consider across development and use.

9. Build an evaluation set before the pilot

Evaluation begins with representative tasks and expected behaviour. Collect questions that reflect real user language, including typos, incomplete details, alternate wording and requests that do not belong in the product. Remove or anonymize personal information in examples unless an approved evaluation process explicitly requires it. For each case, state what a good answer must include, which source should support it and what the assistant should do if the evidence is missing.

Include normal, ambiguous, edge and adversarial scenarios. A normal case tests the everyday flow. An ambiguous case tests whether the assistant asks for the missing detail. An edge case tests a valid but uncommon situation. An adversarial case may attempt to override the assistant’s instructions, extract hidden content or cause it to act outside its scope. The exact test set depends on the system; a small set can start the process, but it should expand as the team observes real failure patterns.

Use a written rubric with observable criteria, such as source relevance, factual support, task completion, clarity, correct refusal or handoff, privacy and appropriate tone. Have a knowledgeable reviewer score examples and note disagreement. A fluent answer can still be wrong, so avoid evaluating only style. Keep a baseline and rerun the same cases after a model, prompt, source index, integration or policy changes. Record the configuration with the result so the team can reproduce what was evaluated.

AI chatbot evaluation cycle: build representative cases, run the production configuration, review with a rubric, improve and retest
Original Heavenkeys evaluation-loop diagram. NIST describes measuring and managing risks in its AI RMF resources; tailor tests to the system, user and impact.

10. Measure the user’s task, not a generic chatbot score

Choose measures tied to the use case. A support assistant might be evaluated on whether users reach the right approved answer, whether they can find the source and whether unresolved requests reach a person. A site guide might be checked for successful navigation to the intended page. An internal drafting assistant might be reviewed for factual corrections and time needed for an employee to complete the task. These are examples, not guaranteed outcomes. Establish the baseline and measurement method before claiming improvement.

Track failures as well as success. Useful categories include unsupported answer, wrong source, missing clarification, inappropriate refusal, privacy concern, broken handoff, incorrect action, confusing wording and latency that interrupts the task. Keep the category definitions clear enough that different reviewers can apply them consistently. If the team combines all cases into a single “accuracy” number, important differences can disappear: a harmless formatting issue and a privacy breach should not carry the same operational weight.

Interpret metrics with their denominator and context. A high response rate may mean the system answers many questions, not that the answers are correct. A low escalation rate may indicate efficient resolution or a failure to recognize uncertainty. A thumbs-up response is feedback from a self-selected subset, not an independent verification of every claim. Review samples, user comments and task outcomes together, and document what each metric cannot tell you.

11. Test prompt injection and other misuse paths

Assume that user input and retrieved content may contain instructions that conflict with the intended behaviour. A document might include malicious or irrelevant text; a user may ask the system to reveal hidden instructions or act on a tool; a conversation may combine benign requests with an attempt to obtain restricted data. Test these paths in a controlled environment using the actual retrieval, access and tool configuration. Do not rely solely on a model’s promise to ignore unsafe instructions.

Review the whole chain: document ingestion, indexing, retrieval filtering, prompt construction, tool permissions, response display and audit logs. Sanitize or constrain data where appropriate, keep secrets out of model context, limit tool capabilities and validate outputs before an action. Consider whether an answer can expose a private URL, internal identifier or source excerpt to the wrong audience. Test both direct and indirect instructions, including content that arrives from an external page or uploaded file.

Record the threat scenarios considered, the controls used and any accepted residual risk. Security testing should be repeated when the system gains new tools or data, not just when the prompt text changes. OWASP’s risk categories can help teams structure an initial review, but the checklist is not exhaustive and does not certify a chatbot as secure. Engage qualified security practitioners for systems involving sensitive data or meaningful business actions.

12. Start with a controlled pilot and explicit stop conditions

Use stages that let the team learn without exposing a broad audience to an untested system. Begin with offline cases; then have internal reviewers examine the same scenarios they will encounter in the service. A limited external pilot can follow only after the owner approves the scope, disclosures, monitoring, escalation route and support plan. Keep a way to pause or disable the assistant without taking down the rest of the website or blocking a human contact route.

Write stop conditions in advance. Examples might include an access-control failure, repeated unsupported policy claims, an unavailable human handoff, a serious data-handling concern or a broken integration that causes users to believe an action succeeded. The exact conditions and thresholds should match the risk and be approved by the accountable owner. A stop rule makes a pilot governable because the team does not have to negotiate from scratch after discovering a serious issue.

Limit what the pilot can do. A read-only assistant with public sources is a different risk from one that updates records or sends messages. If the initial aim is to answer common public questions, do not add transactional tools simply because they are technically available. Collect only feedback needed to evaluate the experience and provide clear notice about the pilot. After the review, decide whether to continue, revise, expand or retire the feature and document why.

Staged AI chatbot rollout from offline tests to internal review, a limited pilot and expansion after agreed checks pass
Original Heavenkeys staged-rollout diagram. It is a planning aid; the organization should set its own approval and stop criteria. See NIST AI risk-management resources.

13. Monitor the service and assign ongoing owners

A launch is not a one-time quality check. Assign an owner for the assistant, its information sources, access controls, provider relationships, user feedback and incident process. Decide who reviews flagged responses, how source changes are approved, who can pause the feature and how the team will communicate a change to users. If these responsibilities sit across teams, name the handoff between them rather than assuming that each team knows the other’s role.

Monitor product quality and service operations together. Quality review may include answer support, citation relevance, correct boundaries, helpful clarification and effective escalation. Operations may include availability, latency, integration failures, provider changes and cost. Safety review may include access events, privacy concerns, misuse patterns and situations where the assistant acted outside its intended role. The right signals depend on the system and should be designed to avoid collecting more user content than necessary.

Set a review rhythm based on risk and change, with immediate review for serious issues and periodic evaluation of routine behaviour. Keep a record of model and configuration changes, source updates, incidents and decisions. When an answer fails, investigate whether the cause was missing content, retrieval, permissions, prompt instructions, a tool, user misunderstanding or a design problem. Fix the underlying issue and add a regression case so the same failure can be checked later.

Ongoing chatbot review covers answer quality, safety, service operation and feedback-driven learning
Original Heavenkeys operations diagram. For an organization-wide risk view, consult the NIST generative AI profile and adapt its guidance to your system.

14. Plan integrations and actions as ordinary software features

A chatbot that reads or changes another system is also an integration project. Document the source of truth, the data exchanged, authentication, error handling, timeout behaviour and ownership of credentials. Decide whether the assistant can only search or can also create, update or delete information. Start with the smallest permission set that supports the task, and verify every action in application code before it is performed.

Design for partial failure. The model may produce a valid-looking request while the CRM, booking system or identity provider is unavailable. The user interface should report what actually happened and offer a safe next step. If a request can be submitted twice, use an idempotent process or confirmation that prevents duplicate actions. If a tool call times out after the remote system has processed it, the product needs a way to check the result before asking the user to try again.

Define what belongs in an audit record and who can review it. The record should help the team investigate an action without unnecessarily retaining full conversations or sensitive data. Retention and access decisions should be approved for the actual business and jurisdiction. Add a human confirmation step for meaningful actions until the team has validated the control and recovery behaviour. A conversational interface should not conceal that a real system is about to be changed.

15. Set expectations for latency, cost and service availability

A useful assistant needs a response time that suits the task, but optimizing for speed alone can encourage incomplete answers or fragile shortcuts. Measure the time a user waits for the first useful response and for the complete task, including retrieval and any external integration. Determine what the interface should show during a longer operation and how it behaves if the provider does not respond. A progress indicator should reflect real work rather than imply certainty.

Estimate costs from the planned usage pattern and configuration, then verify them during the pilot. Include model calls, retrieval infrastructure, logging, evaluation, human review, support and integration operations where applicable. Usage can vary with conversation length, traffic and model choice. Keep a budget owner and set alerts or limits when the service supports them. A cost estimate is a planning input, not a promise of a fixed future bill.

Define how the website behaves if the assistant or model provider becomes unavailable. The public service should still expose ordinary navigation and a human contact route. Avoid placing critical business information only inside chat. If the assistant is optional, it can be disabled while the rest of the site works. If it is part of an essential workflow, define an alternate process and test it before launch.

16. Make the conversation interface accessible and understandable

Give the interface a clear name, purpose and visible entry point, and let users dismiss or ignore it. A chat window should not cover essential navigation, trap keyboard focus or open unexpectedly in a way that interrupts reading. Controls need accessible names and states, and messages should be conveyed to assistive technology without forcing a user to repeatedly search the page. Test keyboard interaction, zoom, contrast, labels and focus movement with the same care used for other interactive features.

Keep responses readable. Break a long explanation into a short answer, a source link and an optional next step. Use headings or bullets when they genuinely help. Avoid presenting a paragraph of generated prose as a verified policy. When the answer contains a limitation, put it near the relevant instruction rather than burying it at the end. Provide a way to correct a misunderstanding and continue through a non-chat route.

Test the interaction with people who use the site in different ways. Automated scans can catch some markup issues, but they cannot tell whether the conversation is understandable or whether a handoff is usable. Ask a participant to complete a realistic task and observe whether the assistant’s controls and output are clear. If a person cannot use the chat, ensure the same service information is available through regular page content or a human contact option.

17. Keep content and product change under governance

Establish an approval process for source material, prompts, evaluation cases and integrations. A product owner can decide what the assistant should do; a content owner can approve the material it answers from; a technical owner can implement the controls; and a privacy or security reviewer can assess the relevant risks. A small organization may combine these roles, but the decisions should remain visible and have an accountable owner.

Track changes that could alter answers: a model version, prompt, retrieval index, permission rule, source document, tool schema or UI disclosure. For a material change, rerun representative evaluations and review the results before expanding the pilot. If a provider makes a change outside the team’s control, monitor the service and have a response path. A system can behave differently because its environment changed even when the project team did not edit its source files.

Retire content and features deliberately. If a policy expires, remove or mark it so that it cannot be retrieved as current. If an assistant is no longer supported, provide a clear alternative, remove credentials and update the relevant pages. Preserve only records the organization has a reason and authority to keep. A feature’s exit plan should be considered alongside its launch plan, especially when user conversations or connected records are involved.

18. Use a planning workshop to define the first release

A focused workshop can turn a broad chatbot idea into a reviewable brief. Invite the business owner, people who handle the task today, a content owner, a technical representative and a reviewer for privacy or security when the use case needs one. Start with a real scenario and map the current steps. Identify where people get stuck, which information is approved and what the team wants the first release to prove.

Leave the workshop with a short set of decisions: intended audience, supported task, excluded requests, approved source list, information handled, access model, escalation route, integrations, evaluation cases, pilot size, stop conditions, operational owner and success evidence. Write down unresolved questions separately, with a person responsible for answering them. Do not hide unanswered policy or privacy decisions inside an implementation ticket.

Use a prototype to review the interaction before building the full system. Show what the assistant says when it has a source, when it needs clarification, when it cannot answer and when a person is needed. Test whether visitors understand the disclosure and can continue without chat. A prototype helps stakeholders discuss the product’s boundaries, not only its visual style. Treat the approved brief as a baseline that can change when evaluation reveals a meaningful issue.

19. A launch-readiness checklist for a business chatbot

Before a pilot, confirm that the task is specific, the simpler alternatives were considered, and the source owner has approved the material. Verify the user disclosure, data flow, retention choices and access controls. Confirm that private records are filtered before retrieval and that tools have narrow permissions. Prepare representative evaluation cases, a written review rubric, human escalation, an operational owner and a tested path to pause the assistant.

For each requirement, name evidence that shows it is met. Examples include a source inventory, an approved scope statement, a data-flow diagram, a permission test result, an evaluation report, a support rota or an incident runbook. A checkbox without evidence can create false confidence. A requirement can also be explicitly out of scope, but the accountable owner should understand and approve that boundary rather than leaving the team to guess.

After the pilot begins, sample results, review user feedback and incidents, retest changes and compare the evidence against the original success criteria. Keep the human route available. Expand only when the team understands the observed failure modes and has controls proportionate to the next audience and action. If the system does not improve the task safely, revise the design or stop the project. Continuing simply because a model integration exists is not a product outcome.

20. Compare building, buying and using a simpler tool

Before committing to a custom chatbot, compare the options against the actual job. A hosted assistant may be enough for public questions if its data handling, controls and integration fit the requirements. A custom application can provide more control over identity, workflow and presentation, but it also creates ongoing responsibilities for engineering, evaluation, monitoring and support. A search improvement, clearer page or structured form may solve the task with less operational overhead.

List requirements that would rule out an option: required identity provider, data residency or retention terms, approved source controls, audit events, accessibility needs, integration support, exportability, language coverage, service availability and exit process. Confirm those requirements in documentation, contract terms and a technical test. A feature page may describe a capability without explaining how it applies to your configuration. Ask the provider for evidence relevant to your use case.

Estimate the full operating cost, not only the subscription or model call. Include implementation, data cleanup, access integration, evaluation time, user support, incident response, provider changes and eventual migration. Determine who owns the assistant’s source content and whether it can be exported. If the vendor changes terms or the product no longer fits, the business should know how to disable the assistant, retain required records and return users to a working alternative.

Make the decision reversible where possible. Start with a narrow pilot, keep the source collection in a portable format and avoid tightly coupling unrelated systems to the chatbot. A small first step can expose operational requirements before a larger contract or custom build is approved. The decision record should explain why the chosen approach fits the task and which assumptions still need validation.

21. Design for language, reading level and regional context

Decide which languages the intended users need and whether the business can support the answers in those languages. A model may produce text in many languages, but that does not establish that the terminology, policy meaning or escalation route is correct. Identify the source version for each language, the person qualified to review it and how changes will remain synchronized. If a translation is unavailable, the assistant should not silently improvise a policy version.

Use everyday language that fits the audience. Service instructions, product terms and abbreviations should be explained. Test questions written by people who use the service, including alternate spellings and regional vocabulary. Include a way to switch language or contact a person if the assistant misunderstands. If the organization operates in a specific region, make sure that hours, currency, units, service area and contact information come from approved current sources rather than being inferred from a model’s general knowledge.

Review translated evaluation cases with fluent human reviewers. A direct translation may change tone, legal meaning or the implied level of certainty. Score whether the assistant preserves the intended answer and boundary, not whether the wording matches English sentence by sentence. Keep feedback channels open for people who cannot complete the task in their preferred language. This work should be planned before the public launch rather than assumed to be solved by selecting a multilingual model.

22. Create failure examples that teach the team what to improve

When a test fails, save a privacy-appropriate record of the scenario, system version, expected behaviour, actual result, source material and reviewer decision. Classify the failure by likely cause: the source is missing, retrieval chose the wrong passage, access filtering failed, the model misunderstood an instruction, the interface concealed a limitation, a tool call was incorrect or a person expected the assistant to do something outside its scope. The classification is a starting hypothesis, not proof; investigate the relevant system before choosing a repair.

For example, if the assistant gives an outdated service hour, adding a stronger instruction may not fix a stale source document. If it reveals information belonging to another user, the first question is whether retrieval or the application permission check was wrong, not whether the model apologized. If it repeatedly sends people to the wrong help page, the issue might be vague labels, duplicate documents or confusing source hierarchy. A useful incident review traces the answer through the whole path.

After the fix, turn the case into a regression test. Rerun it against the updated configuration and check nearby cases that might be affected. A change that improves one example can damage another, particularly when it alters broad instructions or retrieval rules. Keep a short note of the trade-off and the reviewer who accepted the result. Over time, a well-maintained failure set gives the organization more useful evidence than a collection of generic prompt tips.

23. Preserve source references in answers where they help

A source link can help a user inspect a policy, continue reading or check the context for an answer. Decide what makes a citation relevant: it should support the particular claim, be current, resolve for that user and point to an understandable page or section. A link to an entire document may be less useful than a direct heading when that anchor is stable. Do not display a citation merely because retrieval returned a file; the passage should actually support the answer.

Test citation behaviour with questions that combine supported and unsupported parts. The assistant may have evidence for one detail and not another. It should distinguish those parts instead of attaching one source to the whole response. If two approved documents conflict, the experience should follow the documented precedence rule or explain that the information needs review. The content owner should be able to correct a bad source association and test the change.

When a source is internal, verify authorization before showing its title, link or excerpt. Even a filename can reveal confidential context. For public content, link to the canonical page and maintain that page as the source of truth. When the organization cannot provide a stable source reference, the assistant can state the limitation and offer a person to confirm the information. Citation styling should make evidence easier to inspect rather than disguising generated text as a verified official statement.

24. Keep the assistant’s record and change history proportionate

Decide what the organization needs to record for support, evaluation, security and accountability. Possible records include a pseudonymous session identifier, system version, source identifiers, tool result, reviewer decision and incident category. Full conversation text may be useful for investigating some issues but can contain personal or sensitive information. Minimize collection, restrict access, define a retention period and document how a request to delete or correct a record is handled.

Use a change log for prompts, sources, retrieval settings, permissions, connected tools, model configurations and disclosures. Record the reason for a change, approver, date, test cases run and any known limitation. If a vendor changes the available model, note the version or configuration information the provider exposes. This makes it possible to compare an answer across releases and to understand which change might have influenced it.

Be careful when logging hidden prompts, provider responses and user messages. Logs should be protected like the underlying information and should never expose API credentials. Use access controls and redaction where appropriate, while making sure the remaining record is still useful for the intended review. The organization should agree on retention and incident practices with its privacy and security owners rather than copying an arbitrary example from another product.

25. Prepare the support team for the questions the assistant creates

A chatbot can reduce one kind of effort while creating another. People may ask why an answer was given, request a source, report a mistake, ask for a human or need help after an action fails. Prepare the support team with a short explanation of the assistant’s purpose, known limits, escalation path and process for reporting a safety or privacy concern. Staff should know how to identify a conversation reference without requesting unnecessary personal information from the user.

Give support staff authority to correct a source or pause an unsafe flow, or define how they escalate that request to the product owner. If they have no path to report a recurring error, the system may continue to frustrate users while the business sees only isolated complaints. Review support themes with the evaluation owner and turn appropriate cases into tests. Make clear which questions the assistant may answer and which need a qualified employee.

Set an expectation for response times that matches actual staffing. A handoff button that leads to an inbox nobody monitors is not a real handoff. If support is available only during certain hours, say so and provide another option. Review the load created by the pilot and adjust the scope if the team cannot handle the resulting work. Sustainable operations are part of product quality: launch only when the organization can own the experience it is offering.

26. Describe the assistant accurately to customers

Marketing and interface copy should describe what the system actually does. If it searches approved help pages, say that; do not call it an expert or imply that it understands every account situation. Avoid promises such as “always accurate” or “instant resolution.” Explain that generated answers may need review when the task calls for it, and make the human route visible. A clear description helps people decide whether the tool is appropriate for their question.

Review labels when the assistant’s capabilities change. A pilot that only answers from public sources should not inherit copy from a future version that connects to private customer records. A feature name or badge can also imply authority, so test the language with people unfamiliar with the project. If the assistant is unavailable, the site should make that state clear and still show the regular service options.

27. Choose a model and architecture against the task

Model selection should follow the use case, evaluation results, data requirements, expected latency and operating cost. A larger or newer model is not automatically better for every task, and a model comparison based only on a polished demo says little about your sources, users or integrations. Compare candidates on the same evaluation cases and production-like configuration. Include answer support, refusal behaviour, response time, language needs and failure patterns in the review.

Decide where retrieval, rules and ordinary application code are more dependable than generation. A date calculation, eligibility check or permission decision usually needs deterministic logic and authoritative records. A language model can help interpret a question or explain a result, but the application can still compute the value and enforce the rule. This separation reduces ambiguity and makes it clearer which component to investigate when a test fails.

Write down why the chosen approach fits and what would cause the decision to be revisited. For example, the candidate might meet the task’s quality threshold but require a longer response time than the interface can support, or a service may change the data terms. Re-run the comparison when those conditions change. The decision record helps the team avoid choosing technology by novelty and connects architecture to the user outcome it is meant to support.

Frequently asked questions about planning an AI chatbot

Does every business need an AI chatbot?

No. A clear FAQ, search page, form or direct route to staff may solve the problem with less complexity. Consider a chatbot when users have a real task that benefits from conversational help and the business can maintain its sources, controls, evaluation and support.

Can a chatbot use my company documents?

It can be designed to retrieve selected documents, but the organization needs to govern source quality, permissions, retention and provider data handling. Start with approved material and verify that each user can access the retrieved content. Uploading a file does not make it safe or accurate.

How do I stop a chatbot from making things up?

No single prompt can guarantee that every answer is correct. Limit the task, provide approved evidence, require useful source links where appropriate, define uncertainty behaviour, test representative cases and monitor results. For unsupported or high-risk questions, route to a qualified person.

Should the assistant take actions for users?

Only when the business has a clear use case and the application can enforce identity, permission, confirmation, validation and recovery. Begin with read-only assistance when that can demonstrate value. Do not rely on generated text to authorize a sensitive action.

How can I measure whether the chatbot works?

Measure the defined user task, answer support, useful source links, successful escalation, failures and operational cost. Set a baseline and a written rubric before the pilot. A response count or satisfaction click alone does not show correctness or safe task completion.

Can I use customer conversations to train or evaluate it?

That depends on the information, contracts, provider configuration, retention, consent and applicable obligations. Minimize personal data and use approved handling rules. Anonymization may not remove every identifying detail, so involve the appropriate privacy and security reviewers before using real conversations.

What should the chatbot say when it does not know?

It should identify the information it cannot verify, avoid inventing an answer and offer a practical next step. That may be a clarifying question, a link to an authoritative page or a real handoff. The experience should not claim that a person received the issue unless the handoff system confirms it.

How often should chatbot answers be evaluated?

Evaluate before release, after material changes and on a recurring schedule suited to the risk and usage. Review serious failures promptly. Retest when the model, prompt, sources, access rules or integrations change, and add observed failure patterns to the test set.

Is retrieval-augmented generation the same as training a model?

No. A retrieval approach finds selected information at response time and includes relevant material in the model’s context. Fine-tuning changes model behaviour through training. They solve different problems and involve different data, maintenance and evaluation decisions. Choose based on the task and the system’s requirements.

What is the first step if we are not ready to build?

Document the recurring user task, improve the source information and map the current workflow. These steps can reveal that a content or process change is enough. If an AI pilot still appears useful, the team will have a more specific brief and a stronger basis for evaluating it.

Plan for a system people can understand and govern

A responsible chatbot project begins with a bounded task and continues through approved information, access control, evaluation, human support and ongoing ownership. The assistant should make a real task easier while keeping its limits visible. Start with a pilot small enough to review, measure the result against an agreed baseline and expand only when evidence and operating capacity support the next step.

Heavenkeys helps plan and build AI-enabled applications, workflow automation and knowledge search around specific business needs. Explore AI chatbot and voice assistant development, knowledge search and retrieval, and AI evaluation and guardrails. For a related product-planning perspective, read what belongs in a first product release and our customer portal planning guide. Talk with Heavenkeys about a specific task before choosing a model or expanding the scope.