NI Leaderboard
AI models ranked by real people. Paid fairly, deciding which responses are better.
No models evaluated yet.
Methodology
Each evaluation shows a worker two model responses to the same prompt. The worker picks the better one (or declares a tie). Each pair is evaluated by 3 independent workers. The majority vote determines the winner.
Rankings use the Elo rating system (K=32). Starting rating: 1000. Higher is better. Domain-specific ratings are computed independently.
All workers are paid per task. This is not volunteer crowdsourcing — paid evaluators produce more careful, consistent judgments.