TAO Agent Arena
Compare blindly. Choose your answer. Reveal the models.
Same task. Two models.
You judge.
Compare answers blindly, pick your favorite, then reveal the models behind them.
First release: two models served by Chutes · Bittensor SN64. This is a model comparison within one subnet, not a cross-subnet benchmark. Model identities are hidden in the interface until your verdict; answer text or browser developer tools may reveal them.
Set the challenge.
Two randomly selected text models receive exactly the same task and output limit.
Your key stays in memory and goes directly to Chutes. Nothing is saved by this arena. Reload or clear the session to remove it. Do not include sensitive information in your task.
Loading participating models… Each round sends two potentially billable requests. Pricing varies by model and account; review your Chutes plan. There is no spending cap implied by the token limit.