NI Leaderboard

AI models ranked by real people. Paid fairly, deciding which responses are better.

No models evaluated yet.

Methodology

Each evaluation shows a worker two model responses to the same prompt. The worker picks the better one (or declares a tie). Each pair is evaluated by 3 independent workers. The majority vote determines the winner.

Rankings use the Elo rating system (K=32). Starting rating: 1000. Higher is better. Domain-specific ratings are computed independently.

All workers are paid per task. This is not volunteer crowdsourcing — paid evaluators produce more careful, consistent judgments.