DeFAI eval: testing crypto knowledge and exact address extraction
Before an agent touches DeFi, it has to get basic crypto facts right and follow output instructions exactly. defai-eval is a small benchmark we built in February 2025 to check both.
What's in it
The final set has 119 questions from three sources:
- 101 crypto knowledge questions, multiple choice, covering blockchain fundamentals, platform-specific details, trading and finance, crypto history, DeFi, security, layer 2s and more.
- 5 Chromia questions, for example how Chromia's horizontal scaling differs from traditional sharding.
- 13 instruction-following tasks about extraction. The model gets a block of text and must return only the Ethereum, Solana or Tron addresses (or Ethereum transaction hashes) inside it, in an exact format.
The extraction tasks are deliberately noisy. Inputs contain fake hex strings like 0xDEADBEEF, error codes, truncated addresses, repeated addresses and pasted block explorer pages. When nothing valid is present, the model must return exactly ***no_address_found***.
Scoring
Every answer must be wrapped in triple asterisks. The grader pulls out the marked span, strips spaces and brackets, and checks for an exact string match with the ground truth. No partial credit: one missing or extra address scores zero. All models ran at temperature 0.5, and o3-mini ran with high reasoning effort.
Results
Overall accuracy across all 119 questions, from the run logs:
Model | Overall | Instruction following (13) |
|---|---|---|
claude-3-5-sonnet | 92.4% | 76.9% |
deepseek-r1 | 91.6% | 92.3% |
o3-mini-high | 90.8% | 100% |
qwen-2.5-72b | 86.6% | 69.2% |
deepseek-r1-distill-llama-70b | 85.7% | 53.8% |
gpt-4o-mini | 84.9% | 53.8% |
llama-3.2-3b | 62.2% | 53.8% |
What we took from it
Overall knowledge scores bunch together at the top, but the extraction tasks separate the models. The two reasoning models did best on them. o3-mini-high got all 13 right and deepseek-r1 got 12. Claude 3.5 Sonnet had the top overall score but missed three of the 13 extraction tasks. The distilled 70B R1 model scored the same as gpt-4o-mini on extraction, well behind full R1.
For agent builders, the takeaway is that knowing about crypto and handling crypto data are different skills. An agent that mixes up a decoy hex string with a wallet address is a problem in a way a wrong trivia answer is not.
Limitations
This is a small set. Thirteen extraction tasks and five Chromia questions are enough to show differences, not to rank models with confidence. Results come from a single run per model, and the knowledge questions are multiple choice only. The repo also holds early data, such as synthetic honeypot addresses, that is not part of the scored set.
