EVAL Engine

EVAL Engine v1: scoring agent replies and recording them on Chromia

The first version of EVAL Engine, built in January 2025, scored AI agents' replies on X and recorded every score on Chromia. This post covers how it worked: the scoring service, the on-chain record and the SDKs.

Scoring a reply

The engine takes two inputs: an original tweet and an agent's reply. Four LLM judges score the reply from 0 to 100:

  • Truth: how well the reply matches the original tweet's intent.
  • Accuracy: factual correctness.
  • Creativity: originality and how human it reads.
  • Engagement: hooks, calls to action and conversation starters, plus tips for improving them.

The judges were written as DSPy signatures with chain-of-thought, backed by llama-3.1-8b-instant on Groq. Each judge returns a score and a rationale. The final score is the plain average of the four. A fifth step uses all four scores to write a recommended reply, with rules such as staying under tweet length and avoiding emojis and quotes.

The service is FastAPI. It has endpoints for character card evaluation, Virtuals SDK evaluation, score history and tweet suggestions. Requests are checked with an API key and rate limited.

Putting scores on chain

The goal was a critic whose ratings are public and auditable. The flow:

  1. The client signs an evaluate_tweet_request transaction.
  2. The engine co-signs it and submits it to Chromia.
  3. The engine runs the judges.
  4. The engine writes results with update_tweet_scores.
  5. The client gets the result back.

The Rell dapp stores an engine entity (name, description, prefix, address) and a tweet_scores entity holding each request's sub-scores, rationales, recommended reply and evaluation status. Queries let anyone look up scores by request ID, by user address or by engine. Chromia has no gas fees, so writing a score on chain doesn't cost the client a transaction fee.

SDKs

The eval-sdk repo is public and ships the Rell dapp, a TypeScript SDK (eval-engine-sdk on npm) and a Python SDK. The client side is short:

1const txHash = await evalClient.signEvaluateTweetRequest(input, reply);
2const response = await evalClient.submitEvaluateTweetRequest(txHash);

The SDK also includes an EngineClient, so third parties can run their own engine under their own prefix. The public docs on GitBook cover setting up a Chromia account, the JS and Python SDKs, and integrations with Virtuals' G.A.M.E framework (lite and Node.js).

Looking back

This was a narrow first system: one content type, four rubric scores and a small, fast judge model. The parts that carried forward were the method of pairing rubric judges with written rationales, and storing results somewhere they can be audited.