How to Evaluate an AI Model Before Purchasing

Buying an ML model? Learn how to evaluate an AI model before purchasing, with technical checks on quality, performance and real-world suitability.

E@Escrozon
Jul 1, 2026
4 min read
9 views
AI software platform protected by escrow and connected automation tools

Model listings are unusually easy to misrepresent, because the claimed metric and the useful metric are rarely the same number. A seller reporting 97% accuracy has told you almost nothing until you know the class balance of their test set. This guide is about extracting the numbers that matter.

Start With the Metric Trap

Accuracy is misleading on imbalanced data. A fraud-detection model on a dataset that is 99% legitimate transactions reaches 99% accuracy by predicting "legitimate" every time. It is useless and it is accurate.

Ask instead for precision (of the things it flagged, how many were right), recall (of the things it should have flagged, how many did it catch), and F1 where you need one number. For imbalanced problems ask for precision-recall AUC, not ROC-AUC, which flatters minority-class performance.

Demand the confusion matrix. It is four numbers and it makes the trade-off impossible to hide.

Ask what the baseline is. A model beating random guessing is not an achievement. What does a simple logistic regression or a keyword rule achieve on the same data? If the gap is small, you are paying for complexity that buys nothing.

Training Data — the Question Behind Every Other Question

Where did it come from, and was it licensed? A model trained on scraped copyrighted material carries legal exposure that transfers with it. This is the single biggest undisclosed risk in model sales.

How large and how diverse? Ask for the class distribution, not just the row count.

Was the split done properly? The commonest genuine mistake is data leakage — related records appearing in both training and test sets, producing brilliant test scores and poor real-world performance. Ask specifically how the split was made. For time-series data, a random split is a leak by definition; it must be chronological.

Can you see a sample? Even a hundred rows tells you a great deal about labelling quality.

Running Your Own Evaluation

This is the part that decides the purchase, and escrow exists to make it possible before payment.

Test on your own data. The seller's benchmark measures their data. Hold back a labelled set they have never seen and run it yourself.

Reproduce their claimed number first. Run their evaluation on their test set. If you cannot reproduce their headline figure, stop.

Then measure on yours. Expect a drop — distribution shift is normal. A catastrophic drop means the model learned their dataset, not the task.

Measure latency honestly. On your hardware, at your batch size, including preprocessing. Inference time quoted for an A100 is not the number you get on a CPU.

Practical Requirements

Model size and hardware. File size, VRAM required at inference, and whether it runs on CPU at acceptable speed. A model needing 40GB of VRAM has an ongoing cost most buyers do not budget for.

Framework and format. PyTorch, TensorFlow, ONNX. ONNX is the most portable; a checkpoint tied to a specific framework version is the most fragile. Ask which exact versions it was built against.

Quantisation. Is a smaller quantised version available, and what does it cost in accuracy?

Fine-tuning. Is the training pipeline included, or only the weights? A model you cannot retrain has a limited life as your data drifts. This is often the difference between a $2,000 and a $10,000 listing, so establish it before negotiating.

The licence on the base model. A fine-tune of a base model inherits that base model's licence. Some permit commercial use, some do not, some require disclosure. The seller's own licence cannot be broader than the one they built on.

Realistic Pricing

  • Pre-trained NLP model: $500–$5,000
  • Computer vision model: $1,000–$10,000
  • Custom fine-tuned model with training pipeline: $3,000–$20,000
  • Full AI system (model plus application): $5,000–$50,000

Weights alone sit at the bottom of each range. Weights plus training code, documented data provenance, and an evaluation harness sit at the top — and are worth it, because you can maintain them.

Verifying Under Escrow

On Escrozon the seller is paid when you confirm receipt (or automatically 30 days after delivery starts if you raise no dispute), so structure the verification window deliberately:

  1. Receive the weights, code, and documentation in the deal chat.
  2. Reproduce the seller's reported metrics on their test set.
  3. Run inference on your own held-out data.
  4. Measure latency and memory on your target hardware.
  5. Confirm the licence chain — base model through to what you are being sold.
  6. Confirm receipt only when the numbers hold.

If the model does not perform as advertised, open a dispute with your evaluation output attached. Reproducible numbers make these disputes straightforward.

Frequently Asked Questions

The seller will not share training data. Is that reasonable? Sometimes — it may be proprietary or contain personal data. But they should still describe its size, sources, and licensing. Refusing to describe it at all is a warning sign, because it is usually where the legal risk lives.

What accuracy should I expect? There is no universal answer. What matters is performance against a sensible baseline on your data. Always ask what a simple approach achieves on the same problem.

Can I fine-tune a model I bought? Technically usually yes, if you have the weights. Legally it depends on the licence chain, including the base model's terms.

What is data leakage and why does it matter so much? Information from the test set influencing training — near-duplicate records split across both, or a feature that encodes the answer. It produces excellent benchmarks and poor production results, and it is the most common reason a purchased model disappoints.

How do I value a model that works but is undocumented? Lower, and significantly. Undocumented weights with no training pipeline are a dead end once your data shifts. Price it as a fixed asset with a limited life, not as a foundation to build on.

Share: