AI Esquire
Menu
Plans from $497/monthBuy Intake AI
Quality assurance

Your AI intake system is not finished when it goes live.

A polished demonstration proves that an AI intake system can work. It does not prove that the system will work reliably across distressed callers, incomplete facts, background noise, advice requests, unusual names, language changes, integration failures, and the creative disorder of actual legal consumers. Deployment is not the end of evaluation. It is when meaningful evaluation begins.

The demo is the most forgiving environment the system will ever see

Demonstrations are orderly. The caller speaks clearly. The matter fits the intended practice area. The calendar is connected. Nobody starts crying halfway through an answer, asks for a settlement estimate, or gives three spellings of an opposing party's name.

Production is different. Real intake contains ambiguity, urgency, accents, interruptions, missing information, and facts that do not fit the script. A system can sound impressive while producing the wrong operational result. It may complete a fluent conversation but omit the callback number, schedule the wrong appointment, miss an urgent fact, or summarize uncertainty as certainty.

Voice quality is visible. Workflow quality is consequential. A firm should judge the system by the record and next step it creates, not by how human the voice sounds.

Start with the outcome the call was supposed to produce

Every intake path should have a defined objective. For a prospective client, that may be a complete contact record, minimum screening facts, an approved disclosure, and a booked consultation. For an existing client, it may be an authenticated message routed to the responsible team. For an advice request, it may be a respectful refusal to answer combined with a documented escalation.

Without an expected outcome, review becomes subjective. One person thinks the conversation sounded warm. Another thinks it took too long. Neither observation tells the firm whether the intake succeeded.

Build the scorecard before reviewing recordings. The fields will vary by practice, but the core questions are stable.

  • Was the caller's identity and reliable contact information captured accurately?
  • Were required matter facts collected without inventing or overstating details?
  • Were conflicts inputs, urgent triggers, and practice-area criteria routed correctly?
  • Did the system remain within approved administrative boundaries?
  • Was the promised transfer, appointment, alert, or follow-up actually completed?
  • Does the summary distinguish the caller's statements from system inference?

A pleasant conversation with an incomplete record is not a successful intake. It is an error delivered politely.

Review three categories, not one random pile of calls

First, review high-risk interactions. These include advice requests, urgent medical or safety facts, angry or distressed callers, conflict indicators, failed transfers, language switches, uncertain deadlines, and any conversation the system flagged for human attention. The cost of a subtle error can be much greater than the cost of an awkward phrase.

Second, review ordinary calls. A rotating sample of apparently successful intakes can expose quiet failures that exception rules never catch. If the system does not know it omitted a field or misunderstood a name, it will not create its own alert.

Third, review negative outcomes. Examine abandoned calls, appointments that could not be booked, transfers that did not connect, duplicate records, and callers who had to repeat information to staff. These are evidence of friction in the client journey.

The sampling approach should reflect the firm's volume, practice area, and tolerance for error. There is no honest universal percentage that makes a system safe. Review all material exceptions, then sample normal performance often enough to detect patterns before they become routine.

Separate an error, a near miss, and a preference

Not every imperfection deserves the same response. A factual omission that changes routing is an error. A failed escalation caught by staff before harm occurs is a near miss. A greeting that one partner considers too formal may simply be a preference.

Collapsing those categories creates two bad outcomes. Either the firm overreacts to harmless stylistic issues and loses confidence in useful automation, or it treats consequential workflow failures as minor tuning requests. A simple severity model creates discipline.

  • Critical: unauthorized advice, a privacy exposure, a false representation, or failure to escalate immediate danger. Pause or narrow the affected workflow and investigate.
  • Material: incorrect routing, missing required facts, failed scheduling, or a misleading summary. Correct promptly and check for the same pattern elsewhere.
  • Minor: awkward wording, repetition, or a nonconsequential formatting issue. Add it to the improvement queue.
  • Preference: a stylistic choice that does not affect accuracy, boundaries, or completion. Decide whether consistency justifies a change.

The firm needs an incident path before it needs one

NIST's voluntary AI Risk Management Framework organizes risk work around governing, mapping, measuring, and managing. Its Manage function addresses post-deployment monitoring, incident response, recovery, change management, and documented handling of errors and near misses. NIST's March 2026 report on monitoring deployed AI also describes persistent practical challenges, including defining monitoring requirements, obtaining representative performance data, and deciding when intervention is necessary.

These documents are not law and do not provide a legal-intake checklist. Their operating logic is useful: a deployed system requires a feedback and response mechanism.

For a law firm, that mechanism can be modest. Identify who receives the alert, who can review the recording and connected records, who decides whether the workflow remains active, how affected callers or staff are contacted when necessary, and where the incident and correction are documented.

The named owner must have authority, not ceremonial responsibility. If a serious issue appears on Friday evening, someone must be able to disable a route, switch calls to a fallback, or require human review. A governance policy that cannot change production behavior is stationery.

Control changes as carefully as the original launch

AI intake systems change even when the firm does not think of itself as changing them. The script is revised. A new practice area is added. Scheduling rules change. An integration is replaced. A model or voice component is updated. Staff begin relying on a summary field for a purpose it was not designed to serve.

Each material change should have an owner, a reason, a test set, an approval, and a date. Test the normal path, predictable exceptions, and the failure fallback. Then monitor early live interactions under the revised workflow. Otherwise, the firm cannot tell whether a correction solved one problem while creating another.

ABA Formal Opinion 512 addresses lawyers' duties when using generative AI, including competence, confidentiality, communication, and supervision. It does not prescribe a universal intake-audit formula, and jurisdictions may impose different or additional requirements. The narrower practical inference is sound: lawyers cannot outsource accountability merely because a vendor supplies the system.

A workable first 30 days

During the first week, review every exception and enough complete calls to verify each major intake path. Correct material defects before expanding volume. During the next three weeks, continue exception review, rotate through ordinary calls, and compare the system's output with what staff actually needed to act.

At the end of the month, evaluate the evidence: completed contact records, qualification accuracy, consultation bookings, failed transfers, escalations, correction time, caller repetition, and any incidents or near misses. Keep the deployment narrow if the evidence is mixed. Expand only when the workflow is stable and the human review burden is understood.

The objective is not perfection. It is visibility, correction, and a defensible boundary between automated work and professional judgment.

The real test begins after launch

Law firms should be skeptical of two claims: that AI is too risky to use, and that a successful demo proves it is ready to trust. Both positions avoid the harder work of management.

Use the technology where it improves access, consistency, and capacity. Define success. Review what actually happens. Record failures without euphemism. Correct the workflow. Keep a human accountable for the result.

That is not resistance to innovation. It is what serious adoption looks like after the sales call ends.

Sources and further reading

Primary and industry sources used to support this page. External guidance should be reviewed in context and for your jurisdiction.

  1. NIST AI Risk Management Framework 1.0The voluntary framework's Govern, Map, Measure, and Manage functions, including documented response to incidents and errors.
  2. NIST, Challenges to the Monitoring of Deployed AI SystemsA March 2026 analysis of practical monitoring gaps, requirements, data, metrics, and intervention decisions.
  3. NIST AI RMF Manage PlaybookSuggested post-deployment practices for monitoring, incident response, recovery, override, and change management.
  4. ABA Formal Opinion 512ABA guidance on lawyers' duties of competence, confidentiality, communication, and supervision when using generative AI.
Put the framework to work

See Intake AI handle your firm's real workflow.

Bring the intake questions, routing rules, or coverage gap you want to improve. We will demonstrate the system against them.

Call Intake AI now (941) 941-6967Book a 30-minute working session