For the complete documentation index, see llms.txt. This page is also available as Markdown.

Accuracy and evaluation

How GeoPard measures AI Assistant accuracy and validates generated plans.

We measure the assistant's agronomic accuracy on a versioned benchmark. We publish every category score, including areas still improving.

The benchmark

  • 503 expert-level agronomy questions, CCA-style.

  • Five categories: soil and water, crop management, nutrient management, pest management and IPM, precision ag specialty.

  • Versioned and re-run on releases. Current version: v2, July 2026.

Current scores

Category
Score

Precision ag specialty

100%

Soil and water

95.8%

Crop management

93.8%

Nutrient management

92.9%

Pest management and IPM

77.5%

Why the question set is private

Public benchmark questions leak into AI training data. Once that happens, scores stop measuring capability. The question set stays private so future scores remain meaningful. We publish the methodology, all category scores, and sample questions on request.

Validation beyond the benchmark

Benchmark scores measure knowledge. Every generated plan is dry-run validated before creation. Checks cover nutrient balance and removal coefficients against extension recommendations. They cover soil-test critical levels and build-up conventions from your lab. They also cover application rate limits, setback distances, and knowledge base rules. If a check fails, no map is generated.

Your role

The assistant supports professional judgment. It does not replace a certified local agronomist. You review and approve every prescription. Validate recommendations against your crop plan, soil tests, product labels, and local regulations.

Last updated

Was this helpful?