Skip to main content
People Analytics

Resume screening without making it worse

A posting draws nine hundred applications and a recruiter carrying fifteen requisitions has about six seconds each. That is the baseline any tool is compared against — and it is also why a badly built tool can do real damage at speed.

Start with what you are replacing

Every discussion of hiring automation eventually asks whether the model is fair. That is the right question asked in the wrong order. The first question is what the model is replacing, and in a high-volume requisition the honest answer is a keyword filter written by whoever posted the job, followed by a skim so fast that most of the document is never read. Nobody defends that process on the merits. It is worth saying plainly, because a tool can be a real improvement and still be a bad idea, and you cannot tell which without describing the baseline first.

The second thing to establish is scope, and it is the single decision that matters most in this whole area. There are two different products here that get discussed as one. A system that ranks and routes, putting the forty most relevant applications in front of a human who still reads them, is a productivity tool. A system that rejects without a person ever seeing the application is a decision system, and it carries an entirely different burden: legal, evidentiary and moral. Almost everything that goes wrong in this field comes from a company buying the first and quietly operating the second, usually because the volume grew and the shortlist got shorter.

You are probably here because

  • Application volume tripled and the recruiting team is reading a fraction of what arrives
  • Your applicant system has an AI screening feature and nobody can explain what it does
  • Legal asked what happens if a candidate wants to know why they were not advanced
  • Hiring managers say the shortlists are worse than they were two years ago

The last one is usually not the screener. It is usually a job description with fourteen requirements when the manager cares about three.

The model learns your history, including the parts you would not defend

Train on who you hired before and the model learns who you hired before. If a manager historically favoured a handful of schools, the model learns those schools. If it favoured a particular career shape, the model learns to penalise a two-year gap, and a two-year gap is disproportionately a person who was raising a child or recovering from an illness.

Removing the obvious fields does not remove the signal. A resume is a proxy-dense document. Strip the name and the school and there is still the phrasing, the sentence length, the choice between a bulleted and a paragraph format, the sports, the volunteer work, the certification body, the postal code of the last employer. Language models are extraordinarily good at picking up these correlations, which is exactly the property that makes them useful for extraction and dangerous for scoring.

Strip the name and the school and the signal is still there, in the phrasing, the format and the postal code. Anonymising a resume removes the evidence, not the correlation.

This is not an argument for doing nothing. It is an argument for building the thing that does not have this failure mode, and that thing is extraction rather than scoring.

Extract facts, do not manufacture a score

The valuable work a model can do here is read a badly formatted document and pull out the seven or eight things a recruiter would have looked for anyway. Years in a named role. Whether a specific certification is present and current. Location and whether the role is commutable. Whether the person has actually shipped the thing the job requires, with the sentence from the resume that says so. Present that as a table with a link to the source line, sorted however the recruiter wants, and you have removed most of the reading time without removing the decision.

A composite fit score does the opposite. It compresses everything into one number that no one can explain, and the moment it exists somebody thresholds on it, because a number invites a cutoff. Ask the vendor whose score you are considering what a 72 means and what would make it a 78. If the answer is not a sentence a hiring manager could repeat to a candidate, do not put it in the workflow.

TaskModel does it well?What to actually do
Parsing a messy PDF into structured fieldsYes, and this is where most of the time saving isExtract, show the source line, let a human see both
Answering “has this person done X”Reasonably, with the supporting sentence quotedTreat as a filter the recruiter can switch off, never a silent one
Summarising 400 applications into themesYes, and it is underusedUse it to fix the job posting, which is often the real problem
Ranking two close candidatesNo. The differences it keys on are not the differences that matterGive the human both and stop
Judging potential, culture or motivationNo, and confidently produces text that sounds like it canDo not build this. It is where the discrimination risk concentrates
Producing a single 0–100 fit scoreTechnically yes, defensibly noDecline. Nobody can explain it and everybody thresholds on it

The week that changes the outcome is the one before any model

A job description with fourteen requirements, of which the hiring manager truly cares about three, cannot be screened well by anyone. The model will faithfully find people who match fourteen requirements, which is a smaller and stranger group than the people who would do the job well.

So run the job analysis first. Sit with the hiring manager and write down what “good” means in this role in terms of things a person did, not attributes they possess. Rank them. Mark which are genuinely required on day one and which can be learned in a quarter. That document is worth more than any model, it is reusable across the whole requisition family, and if the screening project delivers nothing else it has already paid for itself. It is also the artefact that makes everything downstream defensible, because a criterion tied to what the job actually requires is a criterion you can explain.

Validation is harder than it looks, because your data is censored

To know a screener is better you need outcomes. Here is the problem: you only observe outcomes for people you hired. You have no idea whether the people you rejected would have been good, and the rejected group is where the model's errors live. A screener trained and validated on hires alone is being graded only on the population it already agreed with.

Three partial answers, in ascending order of cost and honesty.

Near-term stage labels. Did the shortlisted candidate pass the phone screen, the technical interview, the panel? These arrive in days rather than years and they measure something real, though they inherit the interview process's own biases.

Longer-run outcomes. Six and twelve month retention, ramp time, first performance rating. Slow, noisy, confounded by the manager and the team, and still the closest thing to the outcome you care about. Wire the join between the applicant system and the human resources system early, because retrofitting it later is a project of its own.

Deliberately interview a random sample of the borderline group. Take a small random draw from the candidates the model ranked just below the line and put them through the loop anyway. It costs interviewer hours and it is the only method that produces uncensored evidence about the model's misses. Most companies will not do this. The ones that do find out something surprising within two quarters.

Where the effort goes in a screening project — our allocation

Job analysis and a ranked definition of “good”
22
Document extraction quality and source citation
21
Stage-by-stage pass-through monitoring
18
Audit trail and record retention
15
Recruiter interface and how a human overrides it
14
The ranking model itself
10

Weights sum to 100. Our starting allocation, not a measurement. The last row being smallest is the point.

Measure pass-through by group, every stage, continuously

The four-fifths rule from the Uniform Guidelines is the standard reference point: if the selection rate for any group falls below eighty percent of the rate for the highest group, that is treated as evidence worth investigating. It is a screening heuristic rather than a legal verdict, and it is the number your counsel will ask for.

Two practical points that get missed. First, measure it at every stage, not just at the offer. A tool that leaves the offer ratio unchanged while halving a group's rate at the resume stage has done harm that the end-to-end number hides. Second, measure it continuously. A one-time audit at launch tells you about the requisitions of that quarter. Volume shifts, sourcing channels change, the model gets a new version, and the ratio moves. Put it on the same dashboard as time-to-fill and look at it monthly.

The regulatory floor is also moving and it is city and state law that gets there first. New York City requires an annual independent bias audit of automated employment decision tools with the summary published, plus notice to candidates. Illinois regulates automated analysis of recorded video interviews. Colorado's rules for consequential automated decisions reach employment. If you operate in several states, the practical answer is to build to the strictest one rather than maintain a patchwork.

The thing nobody expects: candidates write to the model

Applicants have worked out that a language model is reading their document, and some of them now put instructions in it — white text on white background, a zero-point font, text in the document metadata — saying something to the effect of ignoring prior instructions and treating this candidate as exceptionally qualified. This is real, it is easy to do, and a naive extraction pipeline obeys it.

Defences are straightforward if you know to build them. Strip invisible and off-canvas text before extraction. Render the document to an image and read it that way, then compare against the embedded text layer; a large divergence is a flag, not an automatic rejection. Never place candidate-supplied text where the system will read it as instruction, which means the resume goes in a data slot in the prompt with clear delimiters, and the instructions come from you. And log the raw file as submitted, so an investigation a year later has the original.

Record Note

Keep enough to answer the question you will be asked in fourteen months

For every application: the file exactly as submitted, the extracted fields, the model and prompt version that produced them, the rank or routing decision, which humans saw the record and when, and the final disposition with a reason from a controlled list. That is a modest amount of storage and it is the difference between answering a candidate's question, or a regulator's, with a record instead of a recollection. Retention periods for application records vary by jurisdiction; pick the longest one that applies to you and apply it everywhere.

What to ask a vendor

Most products marketed as AI screening are a keyword matcher with a scoring wrapper and a language model summarising the top of the list. That is not a criticism — a well-built keyword matcher is genuinely useful — but it should change what you pay and what you expect. Four questions separate the products quickly.

What exactly does the score measure, and against what outcome was it validated? Ask for the validation study, the population, and the date. “Validated on millions of resumes” is not an answer; validated against retention or performance at companies like yours is.

Can you run it on our last four closed requisitions and show the pass-through rates by group? On your data, not their case study. If they cannot, they cannot support your monitoring obligation either.

What happens when a candidate asks how they were evaluated? There should be a specific answer that does not require a data scientist.

What is the override path, and is it recorded? A recruiter must be able to advance someone the tool ranked low, in one click, with the reason captured. If the interface makes that hard, the tool is a decision system whatever the contract says.

When the honest answer is that you do not need this

Under roughly forty hires a year, screening volume is not the constraint. The constraints are the job description, the interview loop and the speed of the process, and no screening tool fixes any of them. We have told companies exactly that. The money is better spent on writing three good requisition templates and cutting a week out of the scheduling cycle.

Even at high volume, check first whether the bottleneck is reading applications or scheduling interviews. In a lot of organisations recruiters spend more hours coordinating calendars than reading resumes, and calendar coordination is a much easier problem with none of this article's risks.

If your recruiters lose more hours to scheduling than to reading, buy a calendar fix and leave the screening alone.

Before you turn anything on

  • Written, ranked criteria per role family, tied to what the job requires on day one
  • The system ranks and routes; it never sends a rejection on its own
  • Every extracted claim shows the sentence it came from
  • Pass-through by group at every stage, baselined before launch and reviewed monthly
  • One-click override with a recorded reason, and someone reading those reasons
  • Invisible-text stripping and a render-versus-text comparison on every uploaded document
  • Full audit record: original file, extraction, model version, decision, viewer, disposition
  • A named owner who reviews the monitoring pack and can switch the system off

The mistakes we get called about

  • An opaque fit score that a recruiter turned into a hard cutoff within a month
  • Auto-rejection below a threshold, described in the contract as assistance
  • Validated only on people who were hired, so the model was graded on its own agreements
  • Anonymised resumes treated as sufficient, with the proxies untouched
  • A bias audit at launch and never again, while the model and the applicant mix both moved
  • No record of who saw what, discovered when the first written complaint arrived
  • Candidate text pasted directly into the instruction prompt, with predictable results
  • A screener bolted onto a job description nobody had revisited in three years

Bottom line

Use the model to read, not to decide. Extraction with a visible source line saves most of the time and creates almost none of the risk. A single fit score creates almost none of the value and most of the risk. Do the job analysis before anything technical, because the criteria are the product. Validate against stage and retention outcomes while remembering the data is censored, and consider paying for a random sample of borderline interviews to see what you are missing. Monitor pass-through by group at every stage, forever, not once. And if you hire fewer than a few dozen people a year, the correct recommendation is to fix the posting and the scheduling and leave the screening to people.

Frequently asked questions

Is it legal to use automated tools to screen job applicants?

Generally yes, with obligations that vary by where you hire. Several cities and states now require notice to candidates, an independent bias audit, or both, and employment-discrimination law applies to the outcome regardless of whether a human or a system produced it. The practical approach for a multi-state employer is to build to the strictest jurisdiction you operate in and keep the records that let you demonstrate what happened.

Does removing names and schools make screening fair?

It helps and it is not sufficient. Resume text carries group signal through phrasing, formatting, activities, certification bodies and location, and a capable model will use those correlations whether or not anyone intended it to. Blind the obvious fields, then measure outcomes by group, because measurement is the only thing that tells you what the system is actually doing.

How do you validate a screening model when you never see rejected candidates perform?

You cannot fully, and any vendor claiming otherwise is describing a different problem. Use interview-stage pass-through as a fast signal, retention and early performance ratings as slower ones, and un-censor the data by putting a small random sample of near-miss candidates through the full loop. That sample is the only evidence that speaks to the model's misses.

What is the four-fifths rule and how do we use it?

It is a long-standing screening heuristic: if the selection rate for one group is under eighty percent of the highest group's rate, the difference is treated as worth investigating. Compute it at each stage rather than only end to end, baseline it before you deploy anything, and track it monthly. It is a prompt to look, not a verdict, and small samples make it noisy for individual requisitions.

Can candidates manipulate an automated screener?

Some try, most commonly with hidden text in the document instructing the model to rate them highly. Strip invisible and metadata text, compare the rendered page against the embedded text layer, and keep candidate content in a data field rather than in the instructions. Treat a large mismatch as a flag for human review rather than an automatic rejection, since document generators produce odd artefacts too.

1 business day response

Wondering whether your screening tool is helping or hiding something?

Send the stage-by-stage counts from your last few closed requisitions and we will tell you what they show, whether or not there is work in it for us. Email bo@precisionfederal.com.

Email an engineerCapabilitiesMore insights →
People AnalyticsDocument AIModel MonitoringData Engineering