Skip to main content
Staffing & Org Design

Your first data hire: analyst, engineer, or scientist

The three titles get used interchangeably and they are not interchangeable. Each produces something different and each needs something different to already exist. Hire the wrong one and you get a talented person doing work they did not sign up for, badly, alone — and gone in nine months. The diagnostic takes ten minutes.

A note on our incentive We are an engineering firm, so we could tell you to contract this out. We think that is right for some of it and wrong for the most important part — the standing analyst role belongs inside your company, and we say so below. Salary figures here are ranges commonly seen in the US market and vary widely by city and industry; check them against your own recruiting data before budgeting.

What each role actually produces

Forget the titles for a moment and describe the work by its output. That is the only definition that survives contact with a job posting, because the same title means different things at a fifty-person company and at a five-thousand-person one.

The analyst answers questions from data that already exists. Their output is a decision or a number somebody trusts: why churn moved last quarter, which customers are worth calling, what the actual margin is by product line. They need data to be reachable and roughly correct before they can work. Given that, a good analyst is the highest-return person on this list, because they are the only one whose output is directly a business decision.

The data engineer makes data exist, reliably, in one place. Their output is pipelines, a warehouse, freshness, and tests that fail loudly when something upstream changes. Nobody outside the team sees their work directly, which is why it is underhired and undervalued right up until the week the numbers are wrong and nobody can say why.

The analytics engineer sits between them and is the role most companies now need first. They take raw tables and turn them into a small set of clean, documented, tested models that everyone agrees on — one definition of “active customer,” one revenue table, tested nightly. Ten years ago the role barely existed. Now that warehouses are cheap and loading tools are commodity, it is often the highest-value first hire, because most companies' real problem is not missing data but four contradictory versions of it.

The data scientist builds something that makes a repeated decision at volume. Forecasting, scoring, matching, ranking, extraction. Their output is a model plus the evidence that it is better than the thing it replaces. They need clean historical data with outcomes attached, a decision that happens often enough to be worth automating, and a path to production. Missing any of the three, they will spend their year doing analyst work with a scientist's salary and a scientist's frustration.

The diagnostic

Ask these four questions about your own company and answer honestly. They identify the constraint faster than any org chart discussion.

1. How long does it take to answer “what were revenue and retention by segment last month?” If the answer is under an hour and comes from one place, your data foundation is fine and you need an analyst. If it takes two days and involves exporting from three systems and pasting into a spreadsheet, the constraint is engineering. Hiring an analyst into that will produce a person who spends most of their time doing manual exports, which is neither what you hired for nor what they will stay for.

2. If two people compute the same metric, do they get the same number? If not — and this is extremely common — your problem is definitions, not tooling and not modelling. That is analytics engineering: one modelled table, one documented definition, tests that catch drift. It is unglamorous and it is usually the highest-value work available at a company between thirty and three hundred people.

3. Do you have a decision that repeats hundreds of times a week, where historical outcomes were recorded? Both halves matter. A decision that happens twice a month is not worth a model. A decision that happens constantly but where nobody wrote down what happened afterward cannot be learned from. If both halves are true, a data scientist has something to work on. If not, that hire will be idle in the way that looks like being busy.

4. When you already have a good number, does anyone act on it? If the answer is no, no data hire fixes it. That is an operating problem, and adding a dashboard to it produces a better-informed version of the same inaction. This is the question people skip and it is the one that most often explains why the last data hire did not work out.

Hire for what is blocking you, not for the title that sounds most advanced. A data scientist in a company with no reliable tables spends the year building tables, and leaves.
What is true todayThe constraintFirst hire
Numbers take days and involve manual exportsData does not exist in one reachable placeData engineer, or a short contract build then an analyst
Numbers exist but disagree with each otherDefinitions and modellingAnalytics engineer
Numbers are fine, nobody has time to ask questions of themAnalysis capacityAnalyst
A high-frequency decision is made by eye, with history recordedAutomation of judgementData scientist or ML engineer
Good numbers exist and nothing changes because of themOperating discipline, not dataNone — do not hire yet
One bounded build with a clear endCapacity for a project, not a standing roleContract it, then hire the person who runs it after

What it costs

These are broad US ranges for base salary in 2026 as commonly advertised. Major coastal markets run above the top of each band; smaller markets and remote roles often sit below the middle. Add roughly 25 to 35 percent for employer costs and benefits to get to the real number.

RoleTypical base rangeWhat you get for it
Data analyst$80,000–$130,000Questions answered, reporting maintained, decisions supported
Analytics engineer$115,000–$170,000Clean modelled tables, one set of definitions, tests that catch breakage
Data engineer$130,000–$190,000Pipelines and a warehouse that stay up without heroics
Data scientist / ML engineer$150,000–$240,000Models in production, measured against a baseline
Tooling at small scale$1,000–$5,000 per monthWarehouse, loading, orchestration, a reporting layer

The tooling line is worth noticing. At the scale of most companies making a first data hire, the entire platform costs less per month than three days of the person operating it. Spending a long time choosing between two similar warehouses is almost always a worse use of the quarter than picking one and getting the definitions right.

You are probably here because

  • You have budget for one person and three departments want a different one
  • Your last data hire left within a year and you are trying to work out why
  • Reporting takes days and you are not sure if that is a tools problem or a people problem
  • Somebody has proposed hiring a data scientist and you want a second opinion

The four diagnostic questions settle it. The section on why first hires fail is the one to read before writing the job description.

Why the first data hire fails

The pattern is consistent enough to predict, and none of the causes are about the person you hired.

They were hired alone. A first data person has nobody to check their work, nobody to ask when a number looks strange, and nobody who will notice if a query has been quietly wrong for four months. That is a heavy load for a mid-level hire and a lonely one for a senior. If you can only afford one person, buy a few hours a month of outside review from someone who will actually read the code. It is cheap and it is the difference between a hire who grows and a hire who quietly stops trusting themselves.

Nobody owned the questions. The role was created because “we should be more data driven,” and on day one the queue is whatever anyone happens to ask. Six months later the person has produced forty dashboards, none of which changed a decision. Before the hire, write down ten specific questions you would want answered in the first quarter, and name who acts on each one.

The title outranked the problem. Someone senior wanted a data scientist because that sounded like the ambitious hire. The company's actual problem was that revenue could not be reported reliably by segment. The scientist spends the year building tables, does it competently, resents it, and leaves for a role with models in it.

There was no path to production. The model works in a notebook and nobody can put it anywhere. If your first data hire is a scientist, you need at least one engineer — on staff or contracted — who can put things into production, or the output stays a picture of a result.

Access was never granted. The first month goes to waiting for credentials to four systems. It is entirely preventable and it happens constantly. Start the access requests before the start date.

What to have in place before they start

  • One place data lands — a warehouse, even a small and imperfect one
  • Read access to the source systems, granted before day one
  • A named business owner who will act on the answers and can arbitrate priorities
  • Ten written questions you want answered in the first quarter
  • An agreed definition of two or three core metrics, or an explicit mandate to set them
  • A route to production if the role involves models — a person, not a hope
  • Somebody who reviews their work, internal or contracted
  • A budget line for tools, so month two is not a procurement exercise

Hire or contract

The honest split is by the shape of the work, not by cost.

Contract the bounded build. Standing up a warehouse, moving five sources into it, building the first modelled tables, putting one model into production. These have an end. An outside team does them faster because they have done them before, and when it is finished there is nothing left to keep them busy.

Hire the standing role. The analyst, above all. Analysis is 20 percent technique and 80 percent knowing your business — which numbers are unreliable and why, which customer is an exception, what happened in March that explains the spike. An outsider spends months acquiring that and takes it with them when they leave. Do not outsource your analyst. It is the one seat where the context is the job.

A sequence that works well: contract the foundation with a fixed end date, and hire the person who will own it afterward while that work is still running, so their first month is spent learning a system that is being built rather than inheriting one that is finished. The contract team should be writing documentation and a runbook for that named person, and you should say so in the agreement.

How well outside help substitutes for each role — our read

Standing up a warehouse and pipelines
90
First modelled tables and definitions
78
Getting one model into production
74
Reviewing a lone hire's work
66
Ongoing analysis of your business
28
Knowing which numbers to distrust and why
14

Our judgment, and it argues against hiring us for the bottom two rows. Institutional memory is not a service.

Interviewing for the role you actually need

Three exercises separate candidates better than any résumé screen, and all three are cheap to run.

Give them a real messy extract. Twenty minutes with a genuine sample of your data — the duplicated rows, the three date formats, the field that means two things. Watch what they do first. Strong candidates start by asking what the data is supposed to represent and who produces it. Weaker ones start cleaning immediately without knowing what correct looks like.

Ask them to define a metric. “How would you define an active customer here?” The answer should include questions back to you and a set of edge cases: refunds, trials, multi-account organizations, someone who bought once eighteen months ago. If they produce a confident single definition with no questions, they will produce confidently wrong numbers.

Ask about a time the data was wrong. Everyone has one. What you are listening for is how they found out and what they changed afterward. A candidate whose answer includes “so I added a test that would have caught it” is telling you something about how they will work when nobody is watching.

For an analyst specifically, weight communication heavily. The job is half persuasion. A brilliant analysis nobody acts on has the same business value as no analysis.

Mistakes we see

  • Hiring a scientist to fix a plumbing problem, then wondering why they leave
  • One person expected to be all four roles, which produces a burnt-out generalist and no depth
  • The hire reports to whoever had headcount, rather than to whoever acts on the answers
  • Tool selection consumes the first quarter while the definitions stay contradictory
  • No review of their work by anyone, so errors persist quietly for months
  • Dashboards counted as output instead of decisions changed
  • Access requests started after the start date, costing a month of a salary you are paying
  • Hiring before anyone can say what decision would change

Bottom line

Describe the roles by what they produce and the choice mostly makes itself. If getting a number is slow and manual, you have an engineering problem. If two people get different numbers, you have a definitions problem and analytics engineering is the answer. If the numbers are fine and nobody has time to ask questions of them, hire the analyst. Only hire a data scientist when a frequent decision, recorded history, and a route to production all exist at once. And if good numbers already exist and nothing changes because of them, the honest answer is that no data hire will help, and you should spend the money elsewhere.

Frequently asked questions

Can one person cover all of these roles?

For a while, at a small company, with a generalist who genuinely enjoys range. It works for perhaps a year and then breaks in a predictable way: whichever part of the job is least visible gets dropped, usually the engineering, and quality declines quietly. If you are hiring one person, be explicit about which two-thirds of the work is theirs and which third you are deliberately not doing yet.

Do we need a warehouse before hiring anyone?

Something that plays that role, yes, even if it is small and imperfect. Without one place where data lands, the first hire spends their time on manual exports regardless of their title. A basic warehouse with a few sources loaded costs little per month and can be stood up in weeks, which is far less than the salary you would otherwise spend watching someone paste spreadsheets together.

Should the data hire report into technology or into the business?

Analysts do better reporting to whoever acts on the answers, because proximity to the decision is most of their value. Engineers do better near the technology function, where code review and infrastructure live. The failure mode to avoid is an analyst who reports to technology and serves a business unit that has no claim on their time, which produces a queue nobody owns.

Junior or senior for the first hire?

Senior enough to work without supervision, because there will not be any. A first data hire is setting the definitions, choosing the conventions and deciding what good looks like, and those decisions persist for years. A junior hire in that seat is being asked to do a senior job alone. If budget forces a junior, buy a few hours a month of experienced review and make it a standing arrangement rather than an occasional favour.

How do we know if we need anyone at all?

Write down the ten questions you would ask in the first quarter and the decision each one would change. If the list is short, vague, or full of things nobody would act on, wait. Buying a reporting tool and giving two hours a week to a curious person already inside the company is a cheaper first step, and it produces the list of real questions that justifies the hire later.

1 business day response

Not sure which role you are actually missing?

Send the four diagnostic answers and we will tell you plainly which hire we would make first, including when the answer is to wait. Email bo@precisionfederal.com.

Email an engineerCapabilitiesMore insights →
Data TeamsHiringAnalytics EngineeringOrg Design