Start with what the database actually holds
A development database is a twenty-year accumulation of good intentions. Records were merged during one migration and not the next. Some households are one record, some are two, some are three with a deceased spouse still receiving mail. Event attendance lives in a spreadsheet on a shared drive. Volunteer hours are in a separate system nobody has looked at since the person who bought it left. Before any analysis means anything, someone has to know how many distinct people and households are actually in there, and the honest answer is usually a good deal fewer than the record count.

Duplicates are not merely untidy. They understate giving per household, which pushes real major-gift prospects into the mid-level pool where nobody calls them. They inflate the donor count in the board report. They make retention look worse than it is, because last year's donor and this year's donor are the same person under two identifiers. And they produce the mailing that arrives twice at one address, which is the version of this problem your donors actually notice and comment on.
A first pass on matching is not glamorous and it is where the return is. Names with initials, maiden names, nicknames, address changes, deceased flags, business records that are really an individual's company, and the classic pair where one record is the couple and another is the spouse alone. Expect it to change your headline numbers. In a database of any age it is normal to find that several percent of records collapse into others, and it is not unusual for the effect on top-donor rankings to change who the gift officers are assigned to.
You are probably here because
- The board asked for retention and three people produced three different numbers
- Development and finance report totals that have never once agreed
- A long-time supporter got a first-time-donor welcome letter
- Someone is proposing a wealth screening purchase and nobody can say what it would change
None of those four is solved by a model. All four are solved by records that agree with each other.
Donor-advised funds broke the arithmetic
This is the single most common reason a nonprofit's own numbers no longer describe its own donors. When a supporter gives through a sponsoring fund, the payment arrives from the sponsor, not from the person. The check or transfer carries the fund's name. The recommendation letter may name the donor, may name an advisor, may say the gift is anonymous, and may spell the name differently from your record.
If those gifts are recorded as revenue from the sponsor, three things break at once. Your retention rate falls, because a fifteen-year donor appears to have stopped giving. Your donor count falls while revenue holds, which reads like a crisis and is an accounting artefact. And your segmentation sends a lapsed-donor appeal to someone who gave last month.
The fix is soft crediting, and it needs to be applied consistently rather than whenever someone notices. Record the hard credit where the money came from so finance can tie to the ledger, record the soft credit to the individual so fundraising can see the relationship, and make every donor-facing report run on soft credits while every financial report runs on hard credits. Then go back through several years of fund gifts and attribute the ones you can, because your retention history is currently wrong in the same direction every year. The same pattern covers employer matching gifts, family foundations that are really one couple, and gifts from a trust.
Retention is the number, and it needs a definition
Sector studies published each year consistently put overall donor retention somewhere in the low forties as a percentage. First-time donor retention is far worse, commonly in the low to mid twenties, while repeat donors who give a second time retain at rates well above sixty percent. Those figures move a little year to year and vary by cause and by organisation size, but the shape is stable and it is the most useful thing in fundraising data: the second gift is the hard one, and everything after it is comparatively easy.
Which makes the definition worth arguing about once, in writing. Is retention measured on fiscal year or a rolling twelve months? Does a donor who gave in the two prior years and skipped this one count as lapsed or lapsing? Are event ticket buyers donors? Are membership renewals? Does a soft-credited fund gift count, and from which date, the recommendation or the receipt? Every one of those choices moves the number by several points, which is exactly why three people produced three answers for the board.
Write the definitions down, put them in the footer of the report, and never change them silently. If you do change one, restate the prior years. A board that discovers retention improved because the definition changed will not trust the next five reports either.
Where the value sits in a first donor-data project — our allocation
How the value divides in the first year, in our experience. The bottom row is what most proposals are about.
Recency, frequency and amount are hard to beat
Ranking donors by how recently they gave, how often, and how much has been the workhorse of direct response for decades, and it remains difficult to beat by a margin that changes any decision. It is transparent, a development director can explain it to a board in one sentence, it needs no infrastructure, and it degrades gracefully when the data is imperfect.
A gradient-boosted model on the same features plus tenure, channel, appeal history and gift designation will typically add something. Whether it adds enough to matter depends on scale. If you mail forty thousand households, a few points of improvement in response is real money. If you mail three thousand, the difference between the model's list and the recency list is a few dozen names, and the honest recommendation is to spend the money on a better appeal and a phone.
So the test is not whether a model is more accurate. It is whether the list it produces is different enough from the simple list to change what anyone does, and whether that difference is worth the cost of maintaining a model nobody in the building can rebuild. We have advised organisations to stay with a spreadsheet, and we would again.
Where a model does earn its keep: the gift officer's week
The strongest case for real modelling is not the mail file. It is the portfolio. A gift officer can hold a few hundred relationships at most and visits or calls perhaps a dozen people in a week. The question of which dozen is a genuine ranking problem with an expensive wrong answer, and the inputs that matter are behavioural rather than demographic: opened the annual report, attended two events, gave to a restricted programme they never gave to before, increased the recurring amount, replied to a stewardship note, showed up as a volunteer.
That kind of engagement signal is usually scattered across an email platform, an event tool, a volunteer system and the giving history, joined by nobody. Bringing it into one timeline per constituent is a data engineering job, not a machine learning one, and in most organisations it is the piece of work with the highest return. Once the timeline exists, even a simple score built on it is useful, because it surfaces the quiet donor whose behaviour changed and whom nobody would otherwise have called.
Development and finance are allowed to disagree, but only in documented ways
The two functions count different things on purpose. Fundraising counts a multi-year pledge when it is committed; finance recognises revenue on its own schedule. Fundraising soft-credits an individual; finance records the sponsoring fund. Fundraising counts an event's gross; finance nets the costs. None of that is a defect. The defect is when nobody has written down the bridge between the two numbers, so every board meeting spends fifteen minutes on why they differ. Build the reconciliation as a standing report with the differences itemised, and the argument ends permanently.
Send us an export and we will tell you what is really in it.
A constituent export and a gift export covering five years, to contact@precisionfederal.com. You get back an estimate of duplicate households, how much of your revenue arrives through sponsoring funds, and what your retention looks like before and after those two corrections. One business day, no charge. Strip personal details first if you prefer; the shape is enough.
contact@precisionfederal.comWealth screening tells you who is rich, which you knew
Screening services append property values, public securities holdings, philanthropic history and modelled capacity. The data is real and the service can be worth buying. What it cannot do is tell you who cares about your organisation, and capacity without affinity produces a prospect list of wealthy strangers that a gift officer will work through once and quietly abandon.
Two practical cautions. Modelled capacity figures are estimates built on public records with wide error bands, and treating a capacity band as a fact about a person's finances leads to asks that land badly. And the giving history a screen returns is mostly gifts large enough to be published, which skews heavily toward institutions that name buildings, so it says little about the donor who gives four figures quietly every year for two decades.
Use screening to prioritise within a group that has already shown affinity, and treat capacity as one input beside engagement rather than as the ranking itself.
| Question | What it actually needs | Honest verdict |
|---|---|---|
| Who should the gift officer call this week? | A joined engagement timeline per constituent | Worth building. The highest-return work in most shops |
| Who is about to lapse? | Clean recency, plus soft credits applied | Mostly answered by recency once the records are right |
| Which appeal worked? | Appeal codes captured at the gift, and a control group | Answerable only if you set it up beforehand |
| What is our real retention rate? | Deduplication, soft credits, a written definition | Answerable this quarter, and usually a surprise |
| Who has capacity we have not asked? | Screening, applied inside an affinity group | Useful as a filter, dangerous as a ranking |
| What is a donor worth over their lifetime? | Long clean history and stable segments | Directionally useful; precise figures oversell what the data supports |
The line you should not cross
Fundraising analytics touches people who never agreed to be analysed, and a nonprofit's relationship with its supporters is the entire asset. A few positions worth holding firmly.
Never model mortality or health. Planned giving programmes are legitimate and the way to run one is to ask people who have told you they are interested, not to score a database on age and infer the rest. If a donor ever learned they were on a list built that way, the damage would not be recoverable.
Keep service recipients out of fundraising data by default. Many organisations serve people and also raise money, and the two datasets should be joined only with explicit, informed consent. For organisations working in health, housing, immigration, recovery or violence, mere membership in a service list is sensitive information and should be treated that way in the schema, not just in policy.
Write the appeal you would be comfortable explaining. A useful test before any targeted campaign: if this donor asked why they received this specific letter, would the honest answer be reasonable to say out loud? Personalisation that outruns what the donor knows you know is the thing that generates the complaint email, and it costs more than the campaign earns.
What we would not build
- A predictive model on a database that has never been deduplicated
- A lapse model that turns out to be recency with extra steps and a maintenance burden
- A lifetime value figure quoted to two decimals from six years of noisy history
- A dashboard finance cannot tie to the ledger, which will be disbelieved and then ignored
- Any scoring built on inferred health, age-based mortality, or service records
- Channel attribution where no appeal codes were captured at the time of the gift
- A model for an organisation with two thousand donors, where a good list and a phone win
What a first project looks like
Ten weeks, in the order we would do it
Step two needs a person, not just an algorithm. Automatic matching handles the obvious cases and produces a middle band of uncertain pairs that a development associate should review, because they know that two records at one address with different surnames are a married couple and that the business record is the donor's practice. Budget for that review time; it is a few days of somebody's attention and it is what keeps a bad merge out of the database, which is much harder to undo than to prevent.
Step three usually produces the moment that justifies the project. Restating several years of retention with soft credits applied tends to show that the decline everyone has been worried about is smaller than reported, or that it started in a different year than anyone thought. Either way it changes the strategy conversation, and it cost nothing but care.
Before you commission anything
- Definitions of donor, household and retention are written down and dated
- Someone has estimated how many duplicate households exist
- Soft credit rules are documented and applied to history, not just going forward
- You know what share of revenue arrives through sponsoring funds and employers
- Every board number can be tied to the ledger through a written bridge
- Appeal codes are captured at the gift, or you accept that attribution is unavailable
- Any model is compared against plain recency ranking before it is adopted
- Service-recipient data is out of the fundraising schema unless consent is explicit
- Somebody in the organisation can maintain whatever gets built
Bottom line
Donor analytics rewards record-keeping far more than modelling. Deduplicate the households, apply soft credits across the full history so fund and matching gifts land on the people who made them, write the definitions down, and build one engagement timeline per constituent. Those four things will change your retention number, your major-gift assignments and your board conversation. After that, if the mail file is large enough to justify it, a model may add a few useful points. If it is not, the correct advice is to keep the spreadsheet and put the money into the ask itself.
Frequently asked questions
Sector studies published annually put overall retention somewhere in the low forties as a percentage, with first-time donor retention far lower, commonly in the low to mid twenties, and repeat donors retaining well above sixty percent. Compare yourself to those figures only after deduplicating and applying soft credits, because both corrections typically move your number upward by several points and in the same direction every year.
Record the hard credit against the sponsoring organisation so finance can tie to the ledger, and a soft credit to the individual so fundraising sees the relationship. Then run every donor-facing report on soft credits and every financial report on hard credits. Apply the rule retroactively across your history, or your retention trend is wrong for every year in which fund giving grew, which is most of the last decade.
Usually a little better, and often not by enough to matter. The question is whether the model's list differs from the recency-ranked list by enough names to change any decision, and whether someone in the organisation can maintain it. At forty thousand households a few points of lift is real money. At three thousand it is a handful of names and a permanent dependency, and the simple list is the better answer.
It can be, used as a filter within a group that has already shown affinity rather than as a ranking of strangers. Modelled capacity is an estimate from public records with wide error bands, and the giving history it returns is skewed toward gifts large enough to be published. It tells you who could give. Your own engagement data tells you who might want to, and that is the harder and more valuable half.
Only if you set it up in advance. Attribution after the fact is guesswork, because a donor may have seen an email, attended an event and received a letter before giving online with no code attached. Capture appeal codes at the point of the gift, keep a holdout group that receives nothing for at least one campaign a year, and accept that some giving is genuinely unattributable rather than assigning it to whichever channel touched last.
