Skip to main content
Agriculture & Geospatial

AI on agricultural data: yield, imagery, and the liability line

Resolution and revisit numbers that decide whether a model is possible. Yield monitors that lie. Soil polygons that are not measurements. And the moment a recommendation becomes a machine-executable prescription, which is where the legal exposure starts.

The part of agricultural AI that actually breaks

Agricultural AI projects rarely fail on the model. They fail at the joins. A yield map arrives with a third of its points as artifacts. A weather grid cell 4 km across stands in for a rain gauge in a field where the July storm missed half the acres. A soil polygon digitized from a 1970s survey gets treated as a measurement. And somewhere near the end, a lawyer asks who owns the prescription file and whether the platform license lets it be used at all. Every one of those is fixable, and every one of them has to be fixed before the modeling means anything.

Our engineers work these problems the same way we work any regulated data domain: establish what each source physically measures, quantify its error, define the spatial and temporal unit at which a join is honest, and decide up front which outputs are advisory and which will drive a machine. That last decision changes the engineering, not only the paperwork.

Where the effort goes on a working ag-data build

Cleaning and reconciling yield-monitor records
94%
Cloud masking and reflectance harmonization
89%
Field-boundary and management-zone geometry
84%
Weather downscaling to field scale
78%
Spatially blocked validation design
72%
Model fitting and tuning
64%

Relative effort weighting from practitioner reading and build experience; illustrative, not a measured statistic.

Resolution and revisit: the two numbers that decide what is possible

Most imagery arguments end the moment somebody writes down the pixel size and the return interval. A 30 m Landsat pixel is roughly a quarter acre. It cannot see a 30-inch row, a single drowned-out pocket, or a replant strip. A 10 m Sentinel-2 pixel is about a quarter of that and starts to resolve management zones on a 40-acre field. A 3 cm drone pixel resolves individual plants. Each of those supports a different class of question, and no amount of modeling moves a source up a tier.

Revisit is the number people get wrong. Nominal revisit is not usable revisit, because optical sensors do not see through cloud. Sentinel-2 nominally returns every five days with the current constellation. Over the Corn Belt in July, the number of scenes that are genuinely clear over a given field is commonly two to four in a month, and the gaps land exactly when canopy dynamics are changing fastest. Any product promising weekly in-season vigor from free optical imagery has to explain what happens in a ten-day cloudy stretch, and the honest answers are radar, a drone flight, or an interpolation the customer is told about plainly.

SourceResolutionRevisitCost and best use
Sentinel-2 (Copernicus)10 m in four bands, 20 m in six, 60 m in three5 days nominalFree. The in-season workhorse for canopy vigor across whole counties. Red-edge bands at 20 m keep working after NDVI saturates.
Landsat 8 / 9 (USGS)30 m multispectral, 15 m panchromatic, 100 m thermal8 days from the pairFree. The thermal band gives canopy temperature, and the archive runs back to the 1980s for trend work.
HLS (NASA harmonized product)30 m, common grid and reflectance scale2 to 3 days combinedFree. Landsat and Sentinel-2 already cross-calibrated, which removes the single most tedious preprocessing job.
NAIP (USDA)60 cm, four-band including near-infraredTwo to three years by stateFree. Boundaries, tile lines, waterways, tree lines, headlands. Leaf-on and infrequent, so not an in-season sensor.
Sentinel-1 (C-band SAR)about 5 m by 20 m in interferometric wide mode6 to 12 daysFree. Sees through cloud. Crop structure, planting and harvest timing, flood extent, surface-moisture proxies.
Commercial smallsat3 to 5 m daily, 30 to 50 cm on taskingDaily to on-demandPaid. Worth it when a specific calendar date matters more than the budget: a hail claim, a stand assessment, a compliance date.
Drone1 to 5 cm RGB, 6 to 10 cm multispectralWhenever you flyPaid per acre. Plot trials, stand counts, replant decisions, damage documentation, calibration truth for satellite work.

What a drone flight actually gives you, and what it costs

Ground sample distance is arithmetic, not marketing. A 20-megapixel camera on a one-inch sensor with an 8.8 mm lens, flown at the Part 107 ceiling of 400 feet above ground level, produces roughly 3 cm per pixel. Drop to 200 feet and it halves. Multispectral heads such as the common five-band units carry lower-resolution imagers, so the reflectance layers land nearer 8 cm at the same altitude even when the RGB frame is far sharper.

Two things separate a usable multispectral flight from a pretty picture. The first is radiometric calibration: a reflectance panel imaged before and after the flight, plus a downwelling light sensor, or the values drift with the clouds and cannot be compared across dates. The second is flight discipline, meaning fixed altitude, fixed overlap around 75 percent forward and 65 percent side, and a consistent solar window. Without both, year-over-year comparison is not defensible.

Regulation caps the economics. Routine commercial flight runs under 14 CFR Part 107, with Remote ID required under Part 89. Beyond-visual-line-of-sight operation still requires a waiver; the FAA's proposed Part 108 rule for routine BVLOS was published for comment in August 2025 and is not final. Until it is, one operator watching one aircraft is the cost model, and that is why drones cover the acres that justify a per-acre fee while satellites cover everything else.

Yield data is the hardest dataset on the farm

A combine yield monitor is an inference chain, not a scale. Grain flow is estimated from an impact plate or an optical sensor, moisture from a capacitance probe, position from GNSS, and area from header width times ground speed. Logging is typically once per second, so at five miles an hour with a 40-foot head, each point represents roughly 60 square meters of ground, recorded 10 to 15 seconds after the grain actually entered the machine.

That delay is why raw yield maps are full of ghosts. Pass ends produce ramp-up and ramp-down artifacts. Partial header widths on point rows record full-width area against partial-width grain. Speed changes smear the flow signal. Moisture drift shifts a whole field. GNSS offset between the receiver and the header puts the value in the wrong place. In practice a meaningful share of the points in an uncleaned map are artifacts, and a model trained on them learns the combine's operator, not the field.

Cleaning is a defined procedure, and USDA-ARS publishes free software for it in Yield Editor, which implements flow-delay correction, start and end pass filters, velocity and position filters, and moisture standardization. We treat cleaning as a versioned, logged pipeline step with the filter parameters recorded per field and per year, because the same raw file cleaned two different ways produces two different answers, and an agronomist who cannot reproduce the map will not trust the model built on it.

A model trained on an uncleaned yield map learns the combine's operator, not the field.

Soil: map units are not measurements

SSURGO, the NRCS soil survey, is the backbone of nearly every agricultural analysis in the United States, and it is regularly misused. SSURGO delivers map-unit polygons digitized at survey scales between 1:12,000 and 1:63,360, with properties expressed as component percentages within each unit rather than as observations at a point. At 1:12,000 the smallest legible delineation is on the order of an acre and a half. The gridded product, gSSURGO, rasterizes that to 10 m, which makes it look far more precise than the underlying survey ever was.

Used correctly, SSURGO is excellent for what it is: drainage class, texture family, available water capacity, restrictive layers, and a productivity index that explains a large share of the spatial pattern in a yield map. Used incorrectly, it becomes a fake covariate that lets a model appear to learn soil when it is learning polygon boundaries. Real point measurements come from grid or zone sampling, typically at two-and-a-half-acre grids in commercial practice, and from apparent electrical conductivity surveys, which give a dense proxy for texture and water-holding capacity at a fraction of the sampling cost. Continuous 30 m gridded products such as POLARIS help where no sampling exists, with the same caution: they are model output, and their uncertainty travels downstream.

Weather: the 4 km problem

Gridded weather is free, well documented, and coarser than agronomy needs. PRISM and gridMET publish at roughly 4 km, Daymet at 1 km, NOAA's Real-Time Mesoscale Analysis at 2.5 km, and the High-Resolution Rapid Refresh forecast model at 3 km with hourly updates. A 4 km cell covers close to a thousand acres. Convective summer rainfall does not respect that boundary; two fields inside one cell can differ by an inch on the same afternoon.

Temperature transfers across a grid cell far better than precipitation does, which is why growing-degree accumulations built from gridded data hold up while water balance built the same way often does not. Corn growing degree units use the modified formula with a 50 degree Fahrenheit base and an 86 degree cap, and that calculation is stable enough to drive staging models across a county. Water balance needs something closer to the field: an on-farm station, a nearby mesonet or SCAN site, radar-derived quantitative precipitation estimates bias-corrected against gauges, or soil-moisture sensors. When a customer asks for a stress index, the first question we ask is where the precipitation is coming from, because that single choice sets the error floor for everything downstream.

The line between a recommendation and a prescription

This distinction is the most consequential one in the entire field, and it is more legal than technical. A recommendation is information delivered to a person who decides. A prescription is a machine-readable file, a shapefile or an ISOBUS task file under ISO 11783, that goes into a rate controller and meters product onto ground. The moment the output stops being read and starts being executed, the analysis has entered the chain of an application, and applications are regulated.

For pesticides, the label is the law. Under FIFRA it is unlawful to use a registered pesticide in a manner inconsistent with its labeling, at 7 U.S.C. § 136j(a)(2)(G). A variable-rate file that drops a rate below the label minimum or pushes it above the maximum in any part of a field is an illegal application, and the exposure follows the applicator and, in a dispute, the party who authored the file. Jurisdiction matters as well: in California a written recommendation for an agricultural-use pesticide has to come from a licensed Pest Control Adviser, so a software product that emits one is either operating through a licensed adviser or generating a document somebody else has to sign.

Nutrients follow a different path. Federal law does not set field-level nitrogen rates, but NRCS conservation practice standard 590 governs nutrient management plans written under USDA programs, state programs impose binding requirements in nutrient-sensitive watersheds, and operations under a CAFO permit carry Clean Water Act obligations at 40 CFR Part 412. A rate map that conflicts with a filed plan can put a producer out of compliance with a program payment or a permit, and the software will be asked to prove what it recommended and when.

Engineering consequence

Build the prescription path to be reproducible on demand

Any output that can reach a controller gets four things: the exact input snapshot with source dates, the model version identifier, per-zone bounds checked against label or plan limits before the file is written, and the identity of the human who approved it. Store the emitted file itself, not a regeneration. Two years later, in a claim, that record is the entire defense.

Who owns the data

There is no broad federal statute assigning ownership of farm data held by a private platform, which surprises people who assume one exists. What governs is the contract. The 2014 Privacy and Security Principles for Farm Data, signed by major grower organizations and equipment and software companies, set the baseline expectations, and the Ag Data Transparent certification evaluates published agreements against them. Reading a license against that framework takes twenty minutes and answers the questions that matter.

  • Raw versus derived. Growers usually retain raw records. Derived layers, zone maps, and trained models are often claimed by the platform. Say which is which in writing.
  • Aggregation and resale. Can de-identified data be pooled and sold? To whom, at what granularity, and does field geometry survive de-identification? It frequently does.
  • Termination. On exit, what exports, in what format, within how many days, and what is deleted from backups and from models already trained.
  • Sublicensing. Rights granted to affiliates, integration partners, and analytics vendors are where the broadest language usually sits.
  • Change of control. Whether the license survives an acquisition, and whether the grower gets notice and an exit.
  • Trade-secret posture. Farm data can qualify as a trade secret under the Defend Trade Secrets Act, 18 U.S.C. § 1836, only where reasonable secrecy measures exist. A broad platform license undercuts that claim.

Data held by USDA sits under a different rule, and it is a strong one. Section 1619 of the Food, Conservation, and Energy Act of 2008, codified at 7 U.S.C. § 8791, bars USDA from disclosing information, including geospatial information, that a producer or landowner provided to participate in a USDA program. That is why Farm Service Agency Common Land Unit boundaries are not a public download despite being the best field-boundary layer in the country. Any product design that assumes CLU access needs to route through a producer consent path or use derived boundaries instead. When the customer is a federal agency, the contract clauses decide instead: FAR 52.227-14 for rights in data, the DFARS 252.227 series on defense work, and the SBIR data-rights clause on SBIR-funded deliverables.

Free federal data that is genuinely usable

A surprising amount of the required data has no license cost at all. NASS Quick Stats provides county-level yield, acreage, and production through a documented API and is the standard external anchor for validating any county-scale model. The Cropland Data Layer, published annually at 30 m and browsable through CroplandCROS, classifies crop type nationwide with strong accuracy for corn and soybeans and considerably weaker accuracy for minor crops, which matters if the target crop is specialty. NAIP supplies the 60 cm four-band base imagery. Web Soil Survey and gSSURGO cover soils. RMA's cause-of-loss and Summary of Business files describe where and why claims occurred, which is a genuinely underused signal for risk modeling.

On the imagery side, Landsat comes through USGS EarthExplorer, Sentinel through Copernicus, and the harmonized HLS product through NASA Earthdata. All three are also indexed as cloud-optimized GeoTIFFs behind public STAC APIs, so a pipeline can query by field boundary, date range, and cloud fraction and pull only the pixels it needs. That single architectural choice usually turns an imagery budget from a storage problem into a compute line item.

Validation: the failure our team looks for first

The most common defect we find in an existing agricultural model is a validation split done at random over pixels or points. Agricultural data is heavily autocorrelated in space and in time, so a random split puts neighboring points from the same pass in both training and test sets, and the reported accuracy is close to meaningless. Reported R-squared values in the high nineties on field-level yield almost always trace back to this.

Honest evaluation blocks on the unit the model will actually face. Leave-one-field-out answers whether the model transfers to ground it has never seen. Leave-one-year-out answers whether it survives a season with different weather, which is the question that decides whether it is useful in a drought. Leave-one-region-out answers whether an Iowa model can be sold in Kansas. Those numbers come back lower, and they are the numbers a buyer can plan against. We also insist on a calibration check on the residuals, because a yield model that is unbiased on average and badly wrong on the low tail is exactly the model that fails in the year somebody needed it.

What this looks like as a build

The systems we deliver in this space share a shape. A boundary and metadata layer that treats the field-year as the unit of record. Ingest adapters for yield files, sampling results, machine logs, and imagery, each with its own cleaning stage and quality flags that travel with the data. A feature layer that joins imagery, soil, and weather at explicit, documented resolutions. A modeling layer with blocked validation and versioned artifacts. And a decision layer split cleanly in two, advisory outputs on one side and controller-bound prescriptions on the other, the second wrapped in bounds checking, approval, and an immutable record of what was emitted.

Cooperatives, state agencies, and agtech platforms come to us with the same three asks: make the historical data trustworthy, make the model honest about what it does not know, and make the output defensible if it is challenged. Those are engineering problems with known solutions, and they are where our engineers and domain specialists spend their time.

Common questions on the ownership and liability line

If the platform hosts the data, does the grower still own it?

Ownership of raw records is usually retained by the grower under the license, but ownership is less important than the rights granted. A broad, perpetual, sublicensable license to use and aggregate the data leaves the grower nominally owning something the platform can still commercialize. Read the grant clause before the ownership clause.

Can we train a model on customer data and keep the model?

Only if the agreement says so, and it should say so explicitly rather than by implication. Trained weights derived from customer data are the asset most often fought over on termination. Settle it at signature and write down whether the model, the derived layers, or both are covered.

Does anonymizing field data remove the risk?

Not by itself. Field geometry is close to a fingerprint, and a polygon plus a county is usually enough to identify an operation to anyone with local knowledge. If de-identification is part of the promise, it has to cover geometry, not only names and account numbers.

Who is exposed when a variable-rate file is wrong?

It depends on the state, the product, and the contract, and it is rarely only one party. The practical protection is the same in every scenario: bounds-check every zone against label or plan limits before the file is written, require a named human approval, and keep the emitted file with its inputs and model version.

Frequently asked questions

What satellite resolution is good enough for field-level crop analytics?

Sentinel-2 at 10 m resolves management zones on typical row-crop fields and is free. Landsat at 30 m works for county and trend analysis and adds a thermal band. Anything requiring plant-level detail, such as stand counts or small replant areas, needs drone or tasked commercial imagery at sub-meter resolution.

Why do yield maps need cleaning before modeling?

Yield monitors estimate grain flow with a lag of 10 to 15 seconds and compute area from header width times speed, so pass ends, partial widths, speed changes, moisture drift, and GNSS offsets all create artifacts. Uncleaned maps carry a substantial share of bad points. USDA-ARS distributes Yield Editor free for exactly this correction.

Is SSURGO accurate enough to use as a model input?

It is accurate for what it represents: soil map units at survey scale, with properties as component percentages. It is not point data, and its smallest delineations are around an acre and a half. Use it for drainage class, texture family, and water-holding capacity. Use grid sampling or electrical conductivity surveys when the question needs real measurements.

When does an analytics output become a regulated prescription?

When it becomes a machine-executable file that meters product. For pesticides, any rate outside the label is an illegal application under FIFRA at 7 U.S.C. § 136j(a)(2)(G). For nutrients, conflict with a filed 590 plan or a CAFO permit creates program and permit exposure. Bounds checking and recorded human approval belong in the software, not in a policy document.

Why are Farm Service Agency field boundaries not publicly downloadable?

Section 1619 of the 2008 Farm Bill, at 7 U.S.C. § 8791, prohibits USDA from disclosing information, including geospatial information, that producers provide to participate in USDA programs. Common Land Unit boundaries fall under that protection, so any product design that assumes open CLU access needs a producer consent path instead.

1 business day response

Have agricultural data you cannot yet trust?

We build the ingest, cleaning, imagery, and modeling layers behind agricultural analytics for agencies, cooperatives, and agtech platforms, with the prescription path built to be defensible.

CapabilitiesMore insights →Start a conversation
UEI Y2JVCZXT9HP5CAGE 1AYQ0NAICS 541512SAM.GOV ACTIVE