Look for friction, not a feature
A useful starting point for an AI conversation is a task someone already finds difficult, repetitive or slow. It might be finding an answer across internal documents, preparing a first draft from scattered notes, or reviewing a queue of incoming requests. The opportunity is in the work itself, rather than in adding a chatbot to a product.
Describe the workflow before choosing a model. Who starts it? What information do they need? Where do they wait, copy information or ask for help? What does a good result look like? A short conversation with the people doing the work often reveals constraints that are invisible in a demonstration.
Consider a hypothetical support team that prepares replies using product documentation and account notes. Observe several requests from arrival to resolution. The slowest step might be locating the right document, checking a product version or waiting for approval. Each points to a different solution. A drafting tool will do little for a queue blocked by missing decisions. Record a baseline: time spent on a representative task, where someone intervenes and what causes rework. Use those observations to compare possible improvements. Ask the people involved which parts they would trust software to handle and which they would still want to inspect.
Separate the predictable from the ambiguous
Not every automation problem needs AI. When the inputs and rules are clear, ordinary software may be easier to test and maintain. A calculation, an approval threshold or a data transfer usually benefits from explicit logic.
AI becomes more interesting when a task involves interpreting varied language, making information easier to explore, or preparing material for a person to review. Even then, the surrounding workflow should keep predictable operations deterministic. For example, an assistant can draft an explanation while the application calculates the figures and enforces access permissions.
Break the proposed workflow into inputs, interpretation, decisions and actions. Interpreting an incoming message may suit a language model, while retrieving an account and checking an entitlement should follow controlled application logic. This separation helps distinguish a misunderstood request from an incorrect rule or a failed integration. For the support example, a first version could retrieve approved documentation and prepare a suggested response. The specialist decides whether to use it. Account changes and outbound messages remain separate actions with explicit controls. Revisit that boundary only when the trial provides evidence that a broader scope would help and can be operated responsibly.
Define what useful means
Before building, choose a narrow outcome. For an internal search tool, useful might mean helping a support specialist locate the right policy and its source. A fluent response is not enough if the source is outdated, inaccessible to that user or unrelated to the question.
Collect a small, representative set of tasks with the people who know the domain. Include routine requests, incomplete inputs and cases where the correct response is to ask a question or decline to proceed. Agree what a reviewer will check: correctness, source relevance, time spent correcting the result and whether the workflow remains understandable.
Run the same examples through the current workflow and the proposed tool. Have a domain reviewer assess whether outputs omit important qualifications or express more certainty than the evidence supports. Keep unsuccessful examples in the evaluation set. Choose a decision rule before interpreting results: for example, require relevant sources, make unsupported statements visible to reviewers, and measure handling time after corrections are included. These are illustrative criteria, not universal thresholds. The acceptable trade-off depends on the consequences of an error. Include the effort of maintaining the evaluation set when deciding whether the proposed workflow remains worthwhile beyond its initial demonstration.
Make failure part of the design
Ask what happens when the system gets something wrong. Can someone spot the error before it matters? Can the action be reversed? Is there a clear owner for review? An assistant that drafts a response has a different risk profile from one that sends it, changes a record or grants access.
For an initial release, keep the scope small and make review explicit. Show the information used to produce a result where practical. Define a fallback when a source is unavailable or the model cannot produce a usable answer. Keep permissions in the application; do not rely on an instruction in a prompt to enforce them.
Check the information path as carefully as the answer. Identify which systems supply data, which documents each user can access and where prompts and responses are retained. Test with people who have different permissions. A correct answer is still a problem if its source should have been inaccessible to that person. Give the operational owner a way to pause the feature and return to the existing process. Record failures without unnecessarily copying sensitive content. Decide who investigates reports and what must be evaluated again after a model, prompt or document source changes. Treat those responsibilities as part of the release.
Use a short discovery to choose the next step
A useful discovery outcome is a decision, not necessarily a prototype. Write down the workflow, the proposed change, the data required, the evaluation approach and the likely operating responsibilities. Include integration effort and the time people will spend reviewing outputs.
If the opportunity still makes sense, test one bounded workflow. If ordinary automation would solve it more cleanly, choose that. If the required data or ownership is missing, address those gaps first. The aim is to make the next investment more informed, rather than to prove that every problem needs AI.
Turn discovery into a short working brief with one workflow owner, the proposed experience, data sources and evaluation criteria. Name dependencies that could prevent delivery, such as inconsistent documents or missing access controls. Resolve them before treating a prototype as a commitment. At the review point, choose whether to continue, narrow the scope, use conventional automation or stop. Include ongoing support, usage costs and human review in that decision. If the trial succeeds, plan a controlled release with a feedback path. If it fails, record what was learned so the next proposal starts with better information and a clearer understanding of the constraints.