Skip to main content
Insurance & Risk

AI for insurance claims review: automation turns a one-off error into a general business practice

Unfair-claims law does not punish a wrong call. It punishes a wrong call made often enough to look like policy. A carrier or TPA automating claims review is buying frequency, which means the statutory question changes before the technical one does.

The statute does not punish the error. It punishes the pattern.

Read Section 3 of the NAIC Unfair Claims Settlement Practices Act (Model #900) before reading anything about claims automation. An act listed in Section 4 becomes an improper claims practice only if it is "committed flagrantly and in conscious disregard" of the act, or if it "has been committed with such frequency to indicate a general business practice to engage in that type of conduct." California Insurance Code 790.03(h) uses nearly the same construction. Every state's version turns on some form of that clause.

Manual claims operations sit on the safe side of that line largely by accident. A hundred adjusters make a hundred different mistakes. The errors are uncorrelated, they do not form a shape, and an examiner pulling a sample finds noise rather than a practice. Automation removes the accident. One rule, one threshold, one prompt, applied to every file, produces correlated errors by construction. If the rule is wrong it is wrong the same way every time, and the record is not a scatter of individual judgement calls. It is a demonstrated general business practice, documented by the carrier's own system, at whatever volume the system processed.

That is the exposure, and it is not new law. It is existing law meeting a new error structure. The penalty arithmetic follows: Model #900 sets a penalty of not more than $1,000 for each violation with an aggregate cap of $100,000, rising to not more than $25,000 for each violation where the conduct was flagrant and in conscious disregard. Per-violation penalties scale with volume, and so does the remediation cost when a regulator orders a look-back across every file the rule touched.

The conclusion is not "do not automate." The industry has already moved: in the NAIC Big Data and Artificial Intelligence (H) Working Group surveys, 92% of 93 responding health insurers and 88% of 193 responding auto insurers reported that they use, plan to use, or are exploring AI and machine learning. The conclusion is that automation has to be designed around which step it touches, and around what the file can prove afterward.

Where a model can sit in the claim

A claim is not one decision. It is a sequence of roughly six operations carrying very different regulatory weight. Automating intake is not the same kind of act as automating denial, and treating them as one program is how carriers acquire exposure they did not price.

Claim-handling stepWhat an automated system can ownWhat the record still has to prove
Intake and file assemblyIngest the notice of loss, the report, the invoice, the medical record set. Classify, extract fields, index, timestamp receipt.That receipt was acknowledged on time. Model Reg #902 gives an insurer fifteen days from notification to acknowledge a claim, and fifteen days to reply to other pertinent communications.
Completeness checksDetect a missing proof of loss, an unsigned form, an absent operative report. Generate the request and track the clock.That the request was necessary. Model #900 §4(K) names delay caused by requiring both a formal proof of loss and duplicative subsequent verification.
Coverage and policy-term lookupRetrieve the governing form, endorsements, limits, deductibles, exclusions. Surface the provision that applies and its text.That the provision cited is the one relied on. Model Reg #902 §7(A): no insurer shall deny on the grounds of a policy provision, condition or exclusion unless reference to it is in the denial.
Estimation and valuationPrice the repair, the treatment, the loss. Compare against schedules and comparables. Flag outliers both ways.That the number was not systematically low. Model #900 §4(D) and §4(H) both reach settlements offered below what the insured is reasonably entitled to.
Investigation and referralRank files for human review. Assemble the evidence packet. Surface the indicators that drove the rank.That the investigation happened. Model #900 §4(F) reaches refusing to pay without a reasonable investigation, and a score is a referral, not an investigation.
The determination and the noticeDraft it. Assemble the citation set. Check the letter against the file it came from.That a person decided. Model #900 §4(L) requires a prompt, reasonable and accurate explanation of the basis for a denial, and in health lines a licensed clinician owns the medical judgement.

The block below is our reading of how much of each step's working time a system can carry before a qualified human has to own and sign the output. These are editorial judgements from statutory text and published claim-handling standards, not a measurement of any carrier's operation. The last row is low on purpose, and not a comment on model quality.

Automation headroom before a human has to own the output

Intake, document capture, file assembly
94%
Completeness and missing-information checks
89%
Coverage and policy-term retrieval
85%
Estimation and valuation, adjuster-confirmed
77%
Investigation ranking and referral routing
68%
The determination itself, including medical necessity
9%

Editorial weighting from statutory text and published claim-handling standards. Illustrative, not a measured statistic.

Most of the cost in a claims operation sits in the first three rows, and most of the cycle time with it. A system that assembles the file, finds what is missing, and puts the governing policy language in front of an adjuster with the citation attached returns real hours without touching the row that carries the statutory risk. That version of claims automation survives an examination, and it is also the version that pays for itself fastest.

The reconstruction standard is the real design constraint

The single most useful sentence for anyone building claims software sits in Section 4 of the NAIC Unfair Property/Casualty Claims Settlement Practices Model Regulation (#902): "Detailed documentation shall be contained in each claim file in order to permit reconstruction of the insurer's activities relative to each claim." The same section requires claim data be accessible and retrievable for examination, with claim number, line of coverage, date of loss, and date of payment or denial or closure without payment, for the current year and the two preceding years.

Reconstruction is a stronger requirement than logging, and most machine-learning platforms satisfy the weaker one. A log says the system produced output X at time T. Reconstruction means an examiner, two years later, can take the file and answer why. Which policy version governed. Which model version scored it. Which documents were in the file at the moment of decision, and which arrived after. Which threshold applied that week. Which human saw what, and when.

Reconstruction is a stronger requirement than logging, and most machine-learning platforms satisfy the weaker one. A log says the system produced output X at time T. Reconstruction means an examiner, two years later, can take the file and answer why.

Three engineering decisions follow, all cheap on day one and expensive in year three. Version and pin everything that participates in a decision: model weights, prompt text, retrieval index, business rules, policy-form library, fee schedule. A record that names a model by product name rather than immutable version cannot be reconstructed once the vendor ships an update. Snapshot the inputs rather than referencing them, because a record pointing at a document store whose contents changed describes a decision nobody made. And keep the record for the horizon the examination uses, not the one the data platform defaults to. Model #902 names the current year plus two preceding; state law and line of business frequently require longer.

Health claims already have a bright line

In health lines the general fair-claims question has been answered specifically. California SB 1120, signed September 28, 2024, added requirements to Health and Safety Code 1367.01 and Insurance Code 10123.135 governing artificial intelligence, algorithms, and other software tools used in utilization review and utilization management. The tool must base its determination on the enrollee's medical or clinical history, the individual clinical circumstances as presented by the requesting provider, and other relevant clinical information in the record. It must not base a determination solely on a group dataset, must not supplant health care provider decision making, must be open to inspection for audit or compliance review by the department, and its performance, use, and outcomes must be periodically reviewed and revised.

The operative sentence removes the ambiguity: the tool shall not deny, delay, or modify health care services based, in whole or in part, on medical necessity, and a determination of medical necessity shall be made only by a licensed physician or a licensed health care professional competent to evaluate the specific clinical issues. "In whole or in part" closes the workaround where a model produces a recommendation and a reviewer rubber-stamps it at volume.

The federal side moved on parallel timing. Under 42 CFR 422.137, a Medicare Advantage organization's utilization management committee must include a majority of practicing physicians, at least one physician independent and free of conflict, at least one member with expertise in the care of elderly or disabled individuals, and, beginning January 1, 2025, at least one member with expertise in health equity. That committee reviews all utilization management policies including prior authorization at least annually. Under 42 CFR 422.568, the standard timeframe for an organization determination is fourteen calendar days, shortened to seven calendar days effective January 1, 2026 for items and services subject to prior authorization, and a denial notice must state the specific reasons for the denial.

The CMS Interoperability and Prior Authorization Final Rule (CMS-0057-F), published February 8, 2024, applies to Medicare Advantage organizations, state Medicaid and CHIP fee-for-service programs, Medicaid and CHIP managed care entities, and qualified health plan issuers on the federally facilitated exchanges. From January 1, 2026 those payers must send prior authorization decision notices to providers that include a specific reason for denial, and must publicly report prior authorization metrics. From January 1, 2027 they must maintain a Prior Authorization API. Both obligations get harder if the decision behind the metric cannot be explained.

The denial notice is the deliverable

Across every regime that touches claims, the same artifact carries the compliance weight, and it is not the model output. It is the letter.

Property and casualty. Model #900 §4(L) reaches failing, on a denial or compromise offer, to promptly provide a reasonable and accurate explanation of the basis. Model Reg #902 §7(A) requires the claimant be advised of acceptance or denial within twenty-one days of receipt of properly executed proofs of loss, in writing, with any policy provision, condition, or exclusion relied on referenced in it. If more time is needed, §7(B) requires notice within twenty-one days giving the reasons, and a further letter every forty-five days while the investigation stays open.

ERISA plans, which is where most TPAs live. 29 CFR 2560.503-1(g) requires an adverse benefit determination notice to state the specific reason or reasons and reference the specific plan provisions on which the determination is based. For group health plans it also requires either the specific rule, guideline, protocol, or other similar criterion relied on, or a statement that one was relied on together with an offer to provide a copy free of charge on request. That clause deserves a slow read from anyone deploying a scoring model inside a plan's claims process, because a model that functions as an internal criterion is disclosable on request. The timing is tight: seventy-two hours for urgent care, fifteen days pre-service, thirty days post-service, each with one limited extension. Appeal review under §(h) must be conducted by a named fiduciary who is neither the original decider nor that person's subordinate. An automated first pass and an automated appeal review are the same system reviewing itself.

Non-grandfathered group health plans and issuers. 45 CFR 147.136 requires the notice to carry the denial code and its corresponding meaning and a description of the plan's or issuer's standard used in denying the claim. The enforcement hook is the sharpest in the set: if a plan fails to strictly adhere to all the requirements of the internal claims process, the claimant is deemed to have exhausted it and may proceed to external review or to remedies under ERISA section 502(a), with the claim treated as denied on review without the exercise of discretion. A process defect does not just create a compliance finding. It hands the claimant the courthouse.

Build to the letter and the rest follows. If the system can emit a notice naming the provision, stating the standard applied, explaining the basis in language a claimant can act on, and citing the documents relied on, the decision behind it was structured well enough to defend. If the letter can only be produced by a human rewriting the model's output from scratch, the automation never reached the expensive part of the work.

What an examiner asks for, and in roughly what order

The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted by the Executive (EX) Committee and Plenary on December 4, 2023, is the clearest published statement of the document set a department expects. It binds only where a state's department issues it, but Section 4 is worth treating as the checklist regardless, because it describes what a market conduct action looks for.

The bulletin grounds itself in law that already existed, citing the Unfair Trade Practices Act (Model #880), the Unfair Claims Settlement Practices Act (Model #900), the Corporate Governance Annual Disclosure Model Act (#305), the Property and Casualty Model Rating Law (#1780), and the Market Conduct Surveillance Model Law (#693). Its expectation is a written AI Systems Program covering governance, risk management controls, and internal audit, with responsibility vested in senior management accountable to the board, and controls proportionate to the risk of what it calls an Adverse Consumer Outcome.

What is requestedWhat satisfies itWhere teams come up short
The written AIS Program and evidence of its adoptionA dated document, minutes showing adoption, a named accountable executive.A policy that exists but postdates the deployment it governs.
Inventory of models that can produce adverse consumer outcomesA live inventory with owner, purpose, decision touched, version, and risk tier per entry.Rules engines, vendor scores, and spreadsheet models nobody classified as AI and therefore nobody inventoried.
Data lineage, quality, integrity, bias analysis, suitability, currencyDocumented provenance for each input, with the freshness window and bias testing method written down.Training data whose origin cannot be reconstructed, and no defined notion of how stale an input may be.
Validation, testing, and auditing, including model driftA pre-registered test plan, results retained per run, a drift metric with a defined action threshold.Drift monitored on a dashboard nobody owns, with no rule for what happens when it moves.
Third-party diligence, contracts, and auditsDiligence records plus contract terms covering data sourcing, IP, confidentiality, and cooperation with regulators.A vendor attestation used in place of testing. NY DFS Circular Letter No. 7 (2024) says plainly that insurers retain responsibility and cannot rely solely on vendor assurances.
Compliance documentation for the specific model under examinationPer-decision records tying a file to a model version, an input snapshot, and a human sign-off.Aggregate performance reporting offered in place of per-claim reconstruction.

Claims rules trail underwriting rules, and that gap is closing

Read the landscape closely and an asymmetry shows: the specific, testable AI rules written so far mostly govern underwriting and pricing, not claims. New York's Circular Letter No. 7, issued July 11, 2024, covers artificial intelligence systems and external consumer data and information sources in underwriting and pricing. It is a detailed instrument, requiring a three-step discrimination assessment with metrics such as adverse impact ratios and an annual search for a less discriminatory alternative. It does not reach claims handling.

Colorado's SB21-169 is structured so the gap can close on a schedule. It requires a risk management framework, information to the commissioner on external data sources used in developing and implementing algorithms and predictive models, an assessment of the framework's results, and an attestation from the insurer's chief risk officer. The commissioner adopts rules "for specific types of insurance, by insurance practice," a design that walks practice by practice rather than legislating all of them at once.

The NAIC is moving on the oversight side rather than the rule side. Its Big Data and Artificial Intelligence (H) Working Group is piloting an AI Systems Evaluation Tool with twelve participating states as of March 2026, with adoption anticipated at the 2026 Fall National Meeting, and a separate Third-Party Data and Models (H) Working Group is building a framework for vendor-supplied data and models. The direction of travel is toward standardized regulator questions rather than a new substantive standard, which reads as sensible: the standard for claims already exists in Model #900, and what regulators lacked was a consistent way to ask about the machinery behind a pattern.

The planning implication is straightforward. Do not build to a claims-specific AI rule, because there mostly is not one yet. Build to the fair-claims statute in force since 1990 and to the documentation set the AI bulletin already describes, and the specific rule, when it arrives in a state, is a mapping exercise rather than a rebuild.

How we build a claims-review system that can be examined

We work these engagements from the record backward. The order matters: every step below is cheap before the model ships and expensive after.

  • Map the decision, not the workflow. Write down every point where the system's output changes what the claimant receives, and treat only those as regulated decisions.
  • Fix the decision record schema first. Claim identifier, model and prompt version, input snapshot hash, policy form version, threshold in effect, output, human reviewer, timestamp, outcome. Nothing ships until a decision writes one.
  • Put the citation in the output, not in a downstream lookup. Retrieval that names the provision and returns its text at decision time is the difference between a defensible denial and a reconstructed one.
  • Separate the recommendation from the determination in the data model. If those are one field, no amount of policy documentation will show that a person decided.
  • Test on adverse cases before average cases. Hold out the denials, appeals, reversals, and complaint files and measure there first. Accuracy on a population dominated by clean approvals says little about the tail that carries exposure.
  • Instrument the disparity metrics you would be asked for. Run them from the first week rather than the first examination, and keep the results whichever way they come out.
  • Set drift thresholds with a named action. An alert that routes to a dashboard is monitoring. An alert that suspends automated routing and pages an owner is a control.
  • Rehearse the examination. Pull ten closed files at random and reconstruct each decision end to end from the record alone. Whatever cannot be answered is the backlog.

The failure modes worth instrumenting

Silent policy change. A fee schedule updates, a form library gets a new endorsement, a vendor pushes a model revision. Nothing announces it, and the denial rate on one claim category moves four points. This is the most common way an automated operation drifts into a general business practice without anyone deciding anything, and the defense is versioning the inputs rather than watching the outputs. The same discipline covers threshold creep, where a confidence cutoff tuned to clear a backlog quietly moves a population of claims from human review to automatic handling.

Appeal rate as the leading indicator. Of everything a claims operation already measures, the overturn rate on internal appeal most directly predicts a regulatory finding. A rising overturn rate concentrated in one claim type, one geography, or one model version is the pattern an examiner looks for, stated in the carrier's own numbers, and it is worth monitoring at a finer grain than operational reporting usually does.

Explanations generated rather than derived. A language model can write a fluent rationale for a decision it did not make on grounds it did not use. That artifact is worse than no explanation, because it is a written record that misstates the basis. Explanations belong to the decision path: the provision retrieved, the field extracted, the rule applied. The related discipline for document work is covered in deterministic extraction versus generation for records.

Bottom line

Claims automation is worth doing, and the return concentrates in the parts of the claim carrying the least regulatory weight: intake, document handling, completeness, and policy retrieval. The weight sits on the determination and the notice, and there the standard is not accuracy. It is reconstruction, explanation, and a human who owns the call. A carrier or TPA that builds the record first can automate aggressively everywhere else and defend all of it. One that builds the model first will eventually be asked to reconstruct ten thousand decisions from a log that only proves they happened.

Frequently asked questions

Can an AI system deny an insurance claim?

It depends on the line and the state. In California health lines, SB 1120 is explicit that an AI, algorithm, or software tool shall not deny, delay, or modify services based in whole or in part on medical necessity, and that only a licensed physician or comparably qualified licensed professional may make that determination. Outside health lines there is generally no per-se prohibition, but the unfair-claims statute still requires a reasonable investigation and an accurate explanation of the basis. The practical constraint is the same: the system can prepare the determination, and a person has to make it.

What records does a claims AI system have to keep, and for how long?

NAIC Model Regulation #902 requires detailed documentation in each claim file sufficient to permit reconstruction of the insurer's activities relative to each claim, with claim data accessible and retrievable for the current year and the two preceding years. State law and line of business frequently extend that. For an automated system, reconstruction means the model version, the inputs as they existed at decision time, the policy version, the threshold in effect, and the human sign-off, all tied to the claim.

Does a scoring model used in claims have to be disclosed to a claimant?

Under 29 CFR 2560.503-1(g), a group health plan's adverse benefit determination notice must identify any internal rule, guideline, protocol, or similar criterion relied on, or state that one was relied on and offer a free copy on request. A model that functions as such a criterion falls inside that language. 45 CFR 147.136 separately requires the denial code with its meaning and a description of the standard used. Treat disclosability as a design input rather than a legal question raised after deployment.

Are third-party claims models the vendor's compliance problem?

No. NY DFS Circular Letter No. 7 (2024) states that insurers retain responsibility for vendor-supplied tools, must set their own oversight standards, and cannot rely solely on a vendor's assurance of non-discrimination, with contracts expected to include audit rights and cooperation with regulatory inquiries. The NAIC AI model bulletin similarly lists third-party diligence, contracts, and audit results among the documents a department may request. The carrier answers for the model regardless of who trained it.

What is the strongest early warning that an automated claims process has drifted?

The overturn rate on internal appeal, segmented by claim type, model version, and geography. It moves before a complaint pattern forms, and it is the same signal an examiner would build from the outside. Most operations have the data and report it too coarsely to see the movement.

1 business day response

Automating a claim-handling step and need the record to hold up?

We build claims document pipelines, policy-language retrieval, and per-decision audit records designed so an examiner can reconstruct any file two years later.

CapabilitiesMore insights →Start a conversation
UEI Y2JVCZXT9HP5CAGE 1AYQ0NAICS 541512SAM.GOV ACTIVE