The science
Why should anyone trust a crowd?
Because, under the right conditions, crowds have been beating experts for centuries. The conditions are the interesting part. This page explains how Fairgrade creates them, with demonstrations instead of equations. The method itself is described in a working paper by its inventors; what follows is the reasoning behind it, not the mathematics inside it.
The old idea
In 1785, the Marquis de Condorcet proved something strange: if each member of a jury is even slightly more likely to be right than wrong, a majority vote becomes more reliable than any single juror, and the bigger the jury, the more reliable it gets. Financial markets aggregate scattered opinions into prices; game-show audiences famously outperform the phone-a-friend expert. Independent, imperfect judgments, combined, produce accuracy none of the judges possess alone.
Where it breaks
The classical result assumes every judge is equally competent, honestly motivated, and independent. Classrooms, like the internet, violate all three. Some graders know more than others. Some do not try. Some talk to each other. Naive averaging treats a careless score and a careful one as the same evidence, so a handful of bad actors can drag the whole estimate. That is why peer grading, done naively, has a bad reputation, and deserves it.
The distribution is information
This is the example the inventors use to open the paper. Twenty people answer a four-option question and split 5, 3, 10, 2. A simple model says C has a fifty percent chance of being right, since half the room chose it. Common sense, and the empirical wisdom-of-crowds literature, say that understates it badly: ten of twenty is strong evidence, and how strong depends on how competent the voters are.
Assigning sound probabilities to each answer is not a trivial problem, and it is the first thing a theory of collective intelligence has to get right. Play with the assumption and watch certainty move.
Twenty people answered a question with four options. You can see the votes — nothing else.
Computed live from a simple accuracy model. The production Marciano Method goes further: it estimates every voter’s accuracy individually, from their track record — Wisdom in the Crowd.
Disagreement is information too
Collective intelligence flows in two directions at once. The distribution of opinion says how likely each answer is. And differences of opinion are evidence in their own right: about the question, which may be harder or more ambiguous than it looked, and about the judges, since a reviewer who is consistently far from the eventual verdict is telling you something about their credibility.
A simple average discards all of this. Two submissions with the same mean score look identical to it, however differently the reviewers felt.
Two submissions, five reviewers each. Both average exactly 80.
Submission A · reviewers agree
Scores 78, 79, 80, 81, 82
Average
80
95% interval
79–81
Submission B · reviewers disagree
Scores 62, 71, 80, 89, 98
Average
80
95% interval
68–92
The average never moves. The interval does: at this spread the verdict on B is only known to within about ±12 points, against ±1 for A. A simple average would report both as 80 and stop there.
Computed live from a plain confidence interval. In the Marciano Method the same disagreement does more work: it flags the question as contested, and it counts as evidence about each reviewer’s credibility.
Judging the judges
Credibility is surmised the way people already do it.
When you decide how much to trust a piece of advice, you are already running the method informally. You remember who was right before. You give weight to expertise, until it is contradicted. You watch for the friend whose favourite restaurant is always in their own neighbourhood. The Marciano Method makes those three instincts explicit, measurable and mathematical, and estimates all of them from the record rather than from anyone’s say-so.
Judgment
The track record. People whose past advice proved good are trusted more, with an adjustment for those who run consistently harsh or consistently lenient. In a classroom: a reviewer whose scores have tracked the eventual verdict earns weight; one who has not, loses it.
Mastery
Demonstrated expertise. Credentials and prior performance carry weight early, when there is little else to go on, and are discounted if the record contradicts them. A reviewer who does not understand the material cannot reliably judge it, however confident.
Bias
Evidence of an agenda or a tilt. The reviewer who favours friends, or work that resembles their own, or who scores everything high to be done quickly, leaves a pattern. The method looks for it, and weights accordingly.
Credibility is earned, not assumed
The Marciano Method drops the classical assumption of equal judges. It estimates two unknowns at once, the quality of each submission and the credibility of each grader, letting each estimate sharpen the other until they converge. Below is a simplified version of that loop, running live. It is rigged in one direction only: against the bad actors you add. Try to break it.
Dashed line — the grade the instructor gave working alone · scale 55–95
A simplified model of the Marciano Method, running live in your browser. Credibility is earned by agreement with the emerging consensus — and re-earned every time you interfere.
Theory, not only data
Why a theory of collective intelligence at all?
One could skip the theory and rely on heavy data-centric analysis. That approach has two problems. Often there is not enough data to produce quality results without an underlying theory: a class of thirty reviews is not a training set. And a model without a theory offers little explainability, which limits the ability to improve it and makes its verdicts hard to defend to the person receiving them.
The Marciano Method sits between the two failures. Unlike a simple average, it does not treat every judge as equal. Unlike a black-box model, it begins from priors about how human collective intelligence behaves and updates them on the evidence it sees. That is why it works from the first assignment of a new course, why every weight it assigns can be explained, and why it improves as the record grows without ever needing the record to be large.
Objections
Asked and answered.
Will students not game it?
Gaming requires beating your own track record. A grader who scores strategically diverges from the emerging consensus, loses credibility, and with it, influence. Try it above: add colluders and watch their weight collapse. Large coordinated majorities in tiny classes can still distort results, which is why review counts and class-size safeguards exist in the product.
What about lazy graders?
Random or careless scores are the easiest pattern to catch: they disagree with everyone, including each other. Their weight falls toward zero and the diligent majority carries the estimate. The lazy grader's own grade, meanwhile, depends partly on the quality of review work, so diligence is rewarded on both sides.
Students are not experts. Why trust them at all?
No single student needs to be an expert. With a well-designed rubric, each peer is a noisy but informative instrument. The method's job is to extract the signal those instruments share and discount the noise they do not. The instructor's benchmark stays in the loop, and the instructor always finalizes.
Why not simply use machine learning?
Two reasons. Most real evaluations do not come with the volume of history a data-only model needs; a class of thirty reviews is not big data. And a model without a theory of how people judge cannot explain its weights, which means it cannot be corrected or improved. The Marciano Method is grounded in probability theory, works from a single class, and explains every weight.
Is this AI grading in disguise?
No. The trust in Fairgrade comes from credibility-weighted human judgment, mathematics that predates the current AI wave by years. AI-assisted scoring exists as an optional, clearly labelled second opinion beside the peer consensus. It never decides.
What about student privacy?
Reviews are structured and pseudonymous to peers; the platform runs on AWS with a FERPA-ready architecture, and your data stays yours. Credibility scores are internal instruments, not public labels.
Beyond grading
Where the science points next.
The engine does not know it is grading essays. It knows how to estimate the truth from many observers of uneven reliability. That problem appears wherever a hard question meets a crowd, and the list below is where the same mathematics applies. Only the first entry is a product. The rest are directions, stated as such, and each will be earned the way the first was: by being right where the answer can be checked.
Deployed
Peer assessment
Students grade one another's work under an instructor's rubric. The instructor's own grade is the ground truth the method is checked against.
Research direction
Expert panels and peer review
Referees, judges and review committees are crowds of uneven reliability judging the same work. The mathematics is unchanged.
Research direction
Forecasting and risk
Where track record should outweigh title. A method that weights by past accuracy is a method for surfacing the analyst who saw it coming.
Research direction
Talent and performance review
Bias is the known failure of the 360-degree review. A method built to detect and discount it is the natural repair.
Research direction
Public discourse
The disinformation problem is, at bottom, an unweighted crowd. Credentials, record and interest, made visible, let the receiver decide.
Provenance
Ten years in the making.
The research
Two professors at NYU Stern begin a decade of work on credibility-adjusted aggregation: how to extract calibrated truth from crowds of unequal reliability, grounded in probability theory rather than in data alone.
The classrooms
The method is piloted where it was invented: 5,000+ students across 60+ courses at NYU Stern, checked against the instructors' own grades.
The prize
Fairgrade wins the $100,000 Rennert Prize, the grand prize of NYU's $300K Entrepreneurs Challenge (2020).
The platform
The method becomes an enterprise-grade platform: rubric design, smart peer assignment, instructor dashboards, LMS integration, and an optional second opinion. In beta today.
Ready to transform your classroom learning experience?
Start for free — no credit card required. Talk to us when your department is ready.
Are you a student? Join a class →