In a study at Stanford Law School, law professors rated AI-written answers better than those of their colleagues in 75 per cent of comparisons. The result surprised the researchers themselves.
Source: Salinas, A. et al. & Nyarko, J. (2026): Law Professors Prefer AI Over Peer Answers. SSRN, abstract 6849678. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6849678

AI’s work was judged better
AI outperformed experienced law teachers at tutoring students, according to a study led by Stanford Law School. In a blind evaluation, professors chose the AI-produced answer over a colleague’s answer about three times out of four.
The study Law Professors Prefer AI Over Peer Answers was published in the SSRN research database in late May 2026. It was led by Stanford law professor Julian Nyarko, who heads the school’s Legal Innovation through Frontier Technology lab. The author group includes researchers from Yale, New York University, and the University of Chicago, among others.
A blind evaluation in contract law
Sixteen contract law professors from 14 US universities took part. All teach the subject from the same textbook, which made comparing answers fairer.
The professors drew up 40 student questions of the kind first-year students typically ask after a lecture or in office hours. The questions fell into four groups: recalling cases and statutes, recalling doctrine, hypothetical scenarios, and legal policy reflection. Participants wrote their own answers to the questions, then assessed a total of 2,918 anonymised comparisons without knowing whether an answer had been written by an AI or by another professor.
AI won consistently
The result was clear. Professors rated the AI answers better in an average of 75.33 per cent of comparisons, and the win rates of the models tested fell within a narrow band of 75.33–75.92 per cent. Such a small spread suggests this was not a fluke of one particular model. The best AI answers reached the same level as the highest-rated human teacher in the study.
AI answers were also flagged as harmful or misleading less often. Professors flagged only 3.53 per cent of AI answers as pedagogically harmful, against 12.06 per cent for their colleagues’ answers.
The AI model used in the study was Google’s Gemini 2.5 Pro (inside NotebookLM), which is now outdated and no longer available in Google’s AI tools. As of 5 June 2026, the current leading model is Gemini 3.1 Pro.
“We were honestly surprised”
According to Nyarko, the magnitude of the results came as a surprise to the researchers. He emphasises that the questions were not simple and the answers not obvious — many of them required combining complex material and applying legal concepts to new situations.
The researchers chose law precisely because it demands judgement, nuanced reasoning, and the ability to operate in situations open to interpretation, rather than mere memorisation.
Yale law professor and co-author Sarath Sanga framed the starting point this way: in most fields, AI is tested with questions that have one right answer, but in law that is often not the case. Two opposing arguments can both be valid. The researchers wanted to find out whether AI reaches the professional level at which lawyers assess each other’s reasoning. It did.
The researchers do warn against over-interpreting the study. In their view, the finding suggests in principle that AI can strengthen legal education, if it is adopted responsibly.




