JudgeGPT Helped Pakistani Judges Clear Backlogs at $38.50 Return Per Dollar Invested

In Brief

  • Trained judges using JudgeGPT resolved roughly 1,848 more cases per year per district, a 6.3% bump
  • The system returned $38.50 in economic value per dollar invested, versus at least $10 for a control group receiving only a generic seminar
  • Ruling quality held steady or improved; trained judges’ outputs beat controls in 59% of pairwise comparisons

JudgeGPT, an AI assistant built on GPT-4 and RAG over 129,235 documents including 128,292 court rulings and 943 Pakistani laws, materially improved the productivity and quality of judicial output in a large-scale field experiment. The study covered 1,559 judges across 118 trial courts — roughly half of all Pakistani trial-court judges — and found that judges who received targeted training on the system resolved significantly more cases than peers who only attended a generic seminar.

Pakistan’s judiciary is chronically under-resourced. The country has fewer than two judges per 100,000 residents versus twenty-two in the EU and thirty in England and Wales. By the end of 2024, civil and criminal pending cases had reached 2.26 million, with 82% sitting in trial courts. Only 25% of the participating judges had ever used an LLM before the experiment, so researchers from ETH Zurich, the New Economic School and Imperial College London split participants into three groups: JudgeGPT plus training; JudgeGPT plus a generic seminar; and a control group in seminar only.

AI access alone produced negligible change. Judges with untrained access logged roughly twenty times over forty weeks and generated fewer than fifty prompts. The trained subgroup logged roughly sixty times and produced more than two hundred prompts, using the system to parse relevant precedents, draft reasoning paths and cross-reference statutes. After forty weeks, trained judges resolved roughly 1,848 more cases per year per district — a 6.3% increase — while even bottom-quartile judges in that group resolved 616 more cases. The appeal rate per 1,000 cases fell slightly, and ruling quality held steady or improved: trained judges’ outputs beat the control in 59% of pairwise blind evaluations versus 42% for controls.

Why targeted training mattered more than raw AI access

The researchers were explicit that the gains came from instruction, not from the tool alone. The six ninety-minute lectures, delivered by ETH Professor Elliott Ash, taught judges how to frame queries, verify citations and integrate AI output with their own reasoning. Without that scaffolding, judges treated JudgeGPT as a curiosity rather than a workflow. The gap between the trained and untrained JudgeGPT arms is almost entirely explained by adoption and prompting skill.

Costs were kept deliberately low to test scalability. The experiment’s direct spending worked out to roughly $250 per judge over forty weeks, returning $38.50 in economic value per dollar from reduced court delays and faster case finalization. Even under a conservative valuation framework, the return was at least $10 per dollar. For a government with constrained budgets and a massive backlog, that ratio makes a strong case for rolling the tool out nationally, provided the training component is preserved.

Bias audits found no statistically significant increases in gender or religious discrimination in the rulings produced with AI assistance. That is a meaningful finding given the size of the sample and the sensitivity of judicial decisions. The researchers caution that the experiment was run in Pakistan’s trial court system; replication in other legal cultures and higher courts is still needed before assuming the same gains generalize.

What JudgeGPT means for legal AI beyond Pakistan

The result matters because it validates the thesis that retrieval-augmented legal AI can be deployed at civil-service scale under real constraints, not just in lab conditions. For the AI Ethics community, the study provides rare empirical evidence that augmentation — not replacement — improves both throughput and fairness metrics simultaneously. It also shows that domain-specific training beats generic prompting, a lesson that applies to medicine, education and permit processing as much as it does to courts.

ETH Zurich, the New Economic School and Imperial College London plan to publish the full dataset and open-source tooling so other governments can replicate the experiment. If the model holds at national scale, the economic value of clearing millions of pending cases could reach billions of dollars per year across low- and middle-income countries with similar judge-to-population ratios.

Critics will argue that any AI-assisted rulings must still be reviewed for due process, procedural fairness and the right to appeal. The study’s data does not speak to long-term legal doctrine evolution. What it does show is that tool use plus training produces measurable public-sector value quickly, without the dystopian trade-offs that AI doomsayers predict for high-stakes domains.

Source: The Decoder

Leave your vote