The size you are is the whole problem
A carrier or program writing somewhere between twenty and three hundred million in premium sits in an awkward spot. The book is big enough that the loss ratio moves several points a year for reasons nobody can name, and small enough that no individual class, state or hazard group carries enough claims to answer the question. Every well-known result in insurance modeling was demonstrated on a book with a hundred thousand claims a year. You have fifteen hundred. That single fact changes which projects are worth doing and which ones produce a confident answer that is mostly noise.

This is not an argument against modeling. It is an argument for a different order of operations. The projects that pay first at your scale are the ones that do not need a rate filing, do not need a decade of history, and can be turned off on a Wednesday afternoon if the underwriting team hates them.
You are probably here because
- A consultant showed you a lift chart and you cannot tell whether it is real
- Your underwriters are drowning in submissions and binding one in ten of them
- The rating plan has not been re-fit since before the current claims system
- A reinsurer or a broker asked what you are doing about analytics and you did not have an answer you believed
The order below is deliberate. Triage first, because it is fast and reversible. Rating last, because it is slow, filed, and unforgiving of a thin sample.
Start with the funnel, not the rating plan
Nearly every conversation opens with the rate. It is the most regulated, slowest and most data-hungry thing you could pick, and it is where a thin book hurts most. The submission funnel sits right next to it, uses data you already have, and nobody has to approve it.
Look at the actual shape. An independent agency force sends in some number of submissions a week. Underwriters clear them, quote a fraction, and bind a fraction of those. In small commercial, hit ratios in the eight to twenty percent range are ordinary, which means eighty to ninety percent of the underwriting hours in the building are spent on business that never binds. That is the largest single pool of recoverable time in the company, and it does not require a single change to the rating plan.
Three things a triage model can do honestly. It can rank incoming submissions by resemblance to what you bind and keep — not by predicted loss, by predicted fit. It can flag the clearance problem, where the same risk arrives from two agencies under slightly different names and two underwriters work it in parallel; in books we have looked at, duplicates run in the mid single digits as a share of submissions and every one is pure waste. And it can route: this one is standard and can be quoted from the data on the application, that one needs a person.
One thing it must not do is decline. The moment a model removes a risk from the flow, you have made an underwriting decision by algorithm, and you now owe an explanation to your own management, your reinsurers, and possibly your state regulator. Rank and route. Let a human decline.
The data work is most of the work, and none of it is interesting
Policy administration, claims and billing are three systems that were never designed to be joined, and the join is at the policy-term level with effective dates that do not agree. Mid-term endorsements change exposure inside a term. Cancellations create partial terms. A vehicle added in month seven is not the same exposure as one on the policy at inception, and if your exposure base is written vehicle count you have quietly overstated it.
Losses have to be developed. Fitting a model on reported incurred losses is the single most common defect we find, and it is fatal for anything long-tail. Workers compensation and commercial auto liability keep developing for years after the accident date; the most recent two policy years always look excellent because the claims have not matured. A model trained that way learns that recent business is good business, which is a statement about the calendar, not about risk.
Premium has to be on-leveled. If you took rate three times in four years and you model against collected premium, the model learns your rate history. On-level to current rates before anything else touches the data.
Losses usually have to be capped. At your size a single seven-figure claim will dominate any severity fit. Cap at a threshold — a hundred thousand and two hundred fifty thousand are both defensible starting points depending on the line — model the capped layer, and handle the excess with a loaded factor rather than pretending you can fit it. Write the threshold down and never let it drift between analyses.
Expect this stage to take longer than the modeling. On a first engagement it is normal for six to ten of the first sixteen weeks to be spent building a clean policy-term exposure and loss file with nothing clever in it at all. That file is the actual asset. The models are cheap once it exists, and it is worth building even if you never fit a model.
Credibility is a constraint, not an excuse
The classical full-credibility standard for frequency — roughly eleven hundred expected claims for an estimate within five percent at ninety percent confidence — is not a bureaucratic relic. It is a straightforward statement about how much noise sits in a Poisson count, and it is the number that should govern how finely you segment.
Run it against your own book. A commercial auto program with nine thousand power units and a six percent frequency produces on the order of five hundred and forty claims a year. Three clean years is sixteen hundred claims, which is barely enough to speak about the book as a whole. Now split by state, class and fleet size and you are asking twenty claims to settle a rating question. Whatever comes out of that cell will be a number, it will have a confidence interval, and it will be meaningless.
The practical response is not to give up on segmentation. It is to model where the claims are and let structure carry the rest: fit at a level that has data, credibility-weight thin cells toward an external loss cost rather than toward your own empty cell, and be explicit about which relativities in the plan are fitted and which are borrowed. An underwriter can live with "we took the bureau relativity here because we only have twenty claims." What destroys trust is a fitted factor of 1.43 that nobody can defend.
Where the effort goes on a first modeling engagement
Our planning split, not a measurement. The point is the ordering: the last two rows are where the conversation usually starts.
You only see the risks you wrote
Your loss data describes the business you bound. It says nothing about the submissions you declined, and it says nothing about the risks your appetite never saw. That is a censored sample, and it quietly limits every model you will ever fit on it. A model trained on bound business can tell you how a risk inside your appetite will behave. It cannot tell you whether the risks you have been turning away were actually fine.
The fix is cheap and slow: start recording declinations now, with a reason code and the submission data as received. Two years of declination history is worth more than any change of algorithm, and almost nobody has it because the declined submission was never worth a database row to anyone.
The mirror-image mistake is treating quote-not-bound data as a risk signal. A quote you lost tells you about your price against the market on that day. It tells you nothing about the loss potential of the account, and models that conflate the two learn to prefer the business nobody else wanted.
The tree model finds it, the linear model files it
Gradient boosting will usually beat a well-built generalized linear model on holdout lift. In our experience the gap is real but modest — a few points of Gini, not a different business — and it shrinks as the GLM gets the interactions it was missing. That observation suggests the deployment pattern that actually works at a smaller carrier.
Fit the boosted model. Do not deploy it. Read what it found: which interactions it leaned on, where the partial dependence bends, which variables it ignored that your plan pays for. Then take the two or three findings that survive an actuary's scrutiny and a hazard explanation, add them to the GLM as explicit terms, and file the GLM. You keep almost all the lift and you retain the property that matters most in a filed plan, which is that a person can read the factor table and explain any individual quote.
For the target: modeling frequency and severity separately is more work than a single Tweedie fit on pure premium, and it is worth it when you are going to have to explain the plan to someone. Frequency has more claims behind it and is more stable; severity is where the credibility problem bites. Keeping them separate makes it obvious which half of a relativity you actually trust.
A word on filings without drifting into legal advice. Several states now require documented testing of models and external consumer data for unfairly discriminatory outcomes, and Colorado's framework is the most developed of them. Whatever your jurisdictions require, the engineering implication is the same: keep the training data, the fitted object, the variable list and the test results as a versioned artifact tied to the filing, so that a question two years from now can be answered by retrieval rather than by reconstruction.
The override is data, not defeat
A score that an underwriter can override will be overridden, and that is correct. The mistake is treating overrides as leakage to be minimized. They are the highest-quality feedback available and they are free.
Log every one: the score, the action, a reason code from a short controlled list, and a free-text field nobody is required to fill in but everybody does. Six months of that log is a better specification of your business than any requirements document. It will tell you that the model does not know about a particular class of insured, or that it is penalizing a fleet size that your best agent's book is built on, or that one underwriter simply does not trust it.
Watch the rate. An override rate above roughly a third means the model is either wrong or not trusted, and you need to find out which before it goes any further. An override rate near zero is not a triumph; it usually means the underwriting team has stopped reading the file, which is a worse outcome than the model being ignored. And track results on overridden risks separately, because that is the only honest way to settle the argument that will otherwise run for years.
The deployment is somebody else's screen
In small commercial the model does not live in your building. It lives in the moment an agent decides which three carriers to send a submission to, and that decision is made on friction. Every additional question on the application costs you submissions, and the agents will not tell you they stopped; the volume will just drift.
So the variables have a budget. Anything derived from what you already collect is free. Anything from third-party data you already buy is nearly free. A new question the agent has to ask the insured has to earn its place against the submissions it will cost you, and in practice that means one or two new fields, not six.
Two engineering requirements follow. The scoring path has to answer inside a quote screen's patience, which is a second or two, not thirty. And when it cannot answer, the quote must proceed on the existing plan rather than failing. A model that blocks quoting when its service is down has converted a modeling project into an outage, and everyone in the building will remember that longer than they remember the lift chart.
Monitoring is mostly about mix
The dashboards people build after launch watch model accuracy. The thing that actually moves is mix. A new score changes what you write, what you write changes who sends you business, and the fastest-moving risk in the whole system is a distribution partner who has worked out where your new plan is soft.
| Watch | What a move means | Cadence |
|---|---|---|
| Submission volume by producer | One agency tripling right after launch has found a seam. This shows up before any loss data does | Weekly |
| Score distribution of bound business | Shifting into a band you did not intend to grow, usually because of competitive position, not appetite | Monthly |
| Mix by class, state and size | The plan is repricing segments relative to each other; make sure you meant to | Monthly |
| Override rate and reasons | A rising rate on one reason code is a missing variable with a name attached | Monthly |
| Actual vs expected by score decile | The real test, and it needs enough development to be worth reading | Quarterly, honestly annually |
Note the last row. At your claim volume, actual-versus-expected on deciles is not a monthly metric no matter how nice the chart looks. Reading it monthly trains everyone to react to noise, which is how a plan gets tinkered into incoherence.
When you do not need a model
Some honest cases where the answer is no.
The segment is too thin, full stop. A specialty book with a few hundred policies and single-digit annual claim counts will not support a fitted plan. Judgment plus a clean exposure file plus an external loss cost is the correct technology, and adding a model adds only false precision.
Your data is not retrievable yet. If nobody can produce a loss run by class for the last five years within a day, that is the project. It is smaller, cheaper, it is the prerequisite for everything else, and it delivers value on its own — most carriers who build that file find two or three things about their own book that surprise them.
The existing plan has never been refreshed. If your relativities came over from a bureau filing years ago and nothing has been re-examined since, a straightforward refresh on properly developed data will capture a good share of the available lift before anyone fits anything sophisticated.
A single claim decides your year. If your results are driven by severity in the tail, the answer is in reinsurance structure and claims handling, not in a frequency model. Say so out loud early, because a modeling project in that situation will produce a beautiful answer to the wrong question.
A realistic first engagement
Submission triage, first cycle
Shadow mode is the step that gets cut for schedule and the one that decides whether the thing survives. Four weeks of scoring submissions that nobody sees costs almost nothing and answers the only question the underwriting team actually has, which is whether this thing agrees with them on the accounts they are sure about.
The mistakes we get called in to fix
- Trained on reported losses, so the last two years look like the best underwriting in company history
- Fitted relativities in cells with twenty claims, presented with three decimal places
- A boosted model in the quote path that nobody can explain to a regulator or an agent
- Six new application questions, and a submission count that fell nine percent and was blamed on the market
- No override logging, so eighteen months later nobody can say whether the model helped
- Monthly actual-versus-expected on a book that generates enough claims for an annual read
Before you start
- Someone can state the annual claim count for every segment you intend to model
- Development, trend, on-level and capping decisions are written down and owned by a named person
- The exposure and loss file reconciles to the financial statements
- Declinations are being recorded, starting now, whether or not you use them yet
- The model cannot decline anything on its own
- The quote path degrades to the existing plan when scoring is unavailable
- Overrides are logged with reasons from the first live quote
- Data, code, fitted object and test results are versioned together and retrievable in two years
Bottom line
A thin book is not a reason to avoid modeling; it is a reason to sequence it differently. Take the funnel first because it is fast, unfiled and reversible. Spend the real money on the exposure and loss file, because everything downstream is either built on it or built on sand. Respect the credibility arithmetic rather than arguing with it, borrow structure where your own data runs out, and keep the filed model simple enough that a person can defend a single quote. Then log the overrides, watch the mix, and be patient about the loss ratio, because at your size the honest verdict is two or three years out and anyone who promises it sooner is reading noise.
Frequently asked questions
For a segment you intend to price, the classical benchmark is on the order of eleven hundred expected claims for a stable frequency estimate. Most smaller carriers reach that for the book as a whole and nowhere near it for individual cells. That does not block a project; it decides where you fit and where you borrow an external relativity instead.
Both, for different jobs. Use the boosted model as an instrument to find interactions the linear model is missing, then put the ones that survive scrutiny into the GLM and file that. You keep most of the lift and retain the ability to explain any individual quote, which matters in a filed plan and matters to your agents.
Usually submission triage and clearance. It uses data you already have, needs no filing, can be turned off, and attacks the largest recoverable pool of time in the company, since the great majority of underwriting hours go to submissions that never bind.
Slowly, and only with development. Compare actual to expected by score band on matured years, hold the mix constant when you can, and treat anything read inside twelve months as noise. In the meantime, track the leading indicators — mix, producer volume shifts and override reasons — because those move first and are readable now.
Have it rank and route, not decline. An automated declination is an underwriting decision made by software, and you will owe an explanation for it to your own management, your reinsurers and possibly your regulator. Ranking gets you nearly all of the time savings without that exposure.
