verityResearch

Independent research · Pre-registered

The criteria are written before the numbers are read.

verityResearch runs small, rigorous experiments on what fine-tuning and tool access actually buy an AI model. Arm definitions, allowed interpretations and pass/fail thresholds are fixed and frozen first. The results are then read against criteria that already existed.

Est. 2026-08-30 Apache-2.0 verityresearch.dev

Latest result

Tool-call rate

Held-out set, n = 106. Share of prompts answered with a working tool call.

Base model 0%

Fine-tuned — mcAgent v20 98%

0 → 98% Same tools available to both arms.

Method

Pre-registration, in three steps

A result is only as good as the moment its criteria were fixed. Every study published here follows the same sequence, and the order is the whole point.

  1. Define the arms

    Each arm is specified in full — which model, which tools, which prompts — before a single evaluation is run. What counts as a valid interpretation of a response is written down at the same time, so it cannot be widened later to accommodate an answer.

  2. Freeze the pass/fail criteria

    The thresholds that decide the verdict are committed alongside the arm definitions. From this point the study can fail. That possibility is what makes the eventual claim worth anything.

  3. Read the numbers against criteria that already existed

    Only now is a result file opened. The verdict is whatever the frozen criteria say it is. This is a claim about method, not about outcomes — it does not promise that findings will be favorable, only that they were not decided after the fact.

Published result — 01 of 01

/results/mcagent/

Beyond Tool Access — mcAgent v20

Pre-registered Fine-tuning Tool use 2 evaluation sets

Verdict

Fine-tuning earns its claim beyond tool access alone, on both evaluation sets.

Rates by arm and evaluation set. Higher is better in every column. Scroll the table sideways to see all columns.
Arm Tool-call rate Answer rate
Held-out set — n = 106
Base model 0% 48%
Fine-tuned — v20 98% 96%
Adversarial probe — n = 43
Base model 2% 13%
Fine-tuned — v20 83% 86%
Open the full write-up

Criteria frozen before any result file was read

Sharpest finding

The base model essentially never emits a working tool call at all, regardless of phrasing. Tool access alone does not confer tool-use behavior at this model size.
From Beyond Tool Access — mcAgent v20, 2026-09-02. Both arms were given the same tools.

About

A small lab with a fixed procedure

What we publish

Narrow, answerable questions about what fine-tuning and tool access actually buy a model — run end to end, with the criteria set in advance and the code open. Studies appear here when their pre-registered criteria have been resolved, whichever way they resolve.

verityResearch is run by Tony Houston, working with AI collaborators.

Details

License
Apache-2.0
Named
2026-08-30