Most ML model disputes are not about performance. They are about a missing file — a tokenizer, a label map, a preprocessing step — that turns working weights into an artefact nobody can run. This checklist is ordered so the items that most often go missing come first.
1. The Files That Make Weights Usable
Weights alone are rarely enough.
- Model weights —
.pt,.safetensors,.h5, or.onnx. Prefer.safetensorswhere offered; pickle-based formats can execute code on load. - Architecture definition or config — the code or config needed to instantiate the model before loading weights.
- Tokenizer or preprocessor — for NLP, the exact tokenizer used in training. A different tokenizer produces silently wrong output rather than an error.
- Label mapping — for classifiers, which index means which class. Without it your predictions are numbers with no meaning.
- Normalisation values — for vision models, the mean and standard deviation used in training. Wrong values degrade accuracy in a way that looks like a bad model.
Those last three are the usual missing pieces, and their absence is not obvious until inference produces nonsense.
2. Code and Environment
- A working inference script with a runnable example.
- Pinned dependencies —
requirements.txtorenvironment.ymlwith versions. "PyTorch" is not a version, and a two-year-old checkpoint often will not load on current releases. - Setup instructions you can follow without asking questions.
- The training script, if you intend to fine-tune. Weights without a training pipeline have a limited life as your data drifts, and this single item often separates a $2,000 listing from a $10,000 one.
- Hardware requirements — VRAM at inference, and whether CPU inference is viable.
3. Evidence It Works
- Metrics appropriate to the task — precision, recall and F1, not accuracy alone. Accuracy on imbalanced data is close to meaningless.
- The confusion matrix. Four numbers that make the trade-off impossible to hide.
- A named baseline. What does a simple model achieve on the same data? A small gap means you are paying for complexity that buys little.
- How the train/test split was made. Ask directly. For time-series data a random split leaks by construction and produces impressive, meaningless scores.
- Latency and memory, measured on hardware comparable to yours.
4. Legal and Licensing
- The base model's licence. A fine-tune inherits it. The seller cannot grant broader rights than they hold.
- Training-data provenance. Where it came from and whether it was licensed. This is the largest undisclosed risk in model sales and the one buyers ask about least.
- Dependency licences. A GPL component in the inference path has implications for what you distribute.
- Commercial-use confirmation in writing, in the deal chat.
5. Support and Continuity
- What happens if it will not run — is there a support window?
- Are known limitations documented? A seller who lists failure cases is telling you they tested properly.
- Is there any versioning, or is this a one-time artefact?
Running the Checklist Under Escrow
The point of escrow is that this list gets completed before payment. On Escrozon funds are held in escrow while you check the delivery, so work in this order:
- Receive all files in the deal chat.
- Set up the environment from the pinned dependencies alone.
- Run the seller's own example and reproduce their reported metrics on their test set.
- Run inference on your own held-out data.
- Measure latency and memory on your target hardware.
- Confirm the licence chain.
- Only then confirm receipt.
If step 3 fails, stop. Not reproducing the seller's own numbers on the seller's own data means something is missing or misstated, and that is a dispute with clear evidence behind it.
Frequently Asked Questions
Which file format should I prefer?
.safetensors for safety, .onnx for portability. A framework-specific checkpoint tied to an exact library version is the most fragile option.
The seller will not share training data. Reasonable? Sometimes — it may be proprietary or contain personal data. But they should still describe size, sources and licensing. Refusing to describe it at all is where legal risk hides.
How much does a training pipeline add to the price? Often several times the value of weights alone, and it is usually worth it. Without it you cannot retrain as your data shifts, so you are buying a depreciating asset.
What if performance is good on their data and poor on mine? That is distribution shift, and some drop is normal. A catastrophic drop suggests the model learned their dataset rather than the task — a legitimate basis for a dispute if the listing claimed general performance.
Do I need an ML engineer to run this checklist? For items 1 and 2, no — you are checking files exist and instructions work. For item 3 onward, yes. Budget a few hours of expert time on any significant purchase; it fits inside the escrow window and is cheap insurance.



