2026 Clause-Extraction Accuracy Benchmark: Sirion vs Open-Source LLMs

Subscribe to our Newsletter

A man in a blue blazer sits on a desk, looking at a tablet in an office enviroment.
  • Sirion’s Extraction Agent achieved a 94.2% overall F1-score in the ContractEval benchmark.
    It led across commercial terms, risk and compliance, and operational clause extraction.
  • Sirion outperformed fine-tuned GPT-4 by 8.9 percentage points on overall extraction accuracy.
    Sirion scored 94.2% compared with 85.3% for GPT-4, 83.9% for Claude 3.5 Sonnet, and 79.2% for Llama 3.1 (70B).
  • Commercial terms delivered Sirion’s strongest extraction performance at 96.1%.
    High extraction accuracy helps enterprises reliably capture critical information such as payment schedules, pricing terms, discounts, and currencies.
  • Accuracy gains did not come at the expense of processing speed.
    Sirion processed contracts in 2.3 minutes on average, compared with 3.9–5.1 minutes for the other models evaluated.
  • Enterprise buyers should validate extraction performance through controlled pilots, not vendor accuracy claims alone.
    F1-score, error rates, processing speed, explainability, scalability, and performance on representative contracts provide a stronger basis for CLM evaluation.
About the author
A man in a blue blazer sits on a desk, looking at a tablet in an office enviroment.

Sirion

Sirion is the world’s leading AI-native CLM platform, pioneering the application of Agentic AI to help enterprises transform the way they store, create, and manage contracts. The platform’s extraction, conversational search, and AI-enhanced negotiation capabilities have revolutionized contracting across enterprise teams – from legal and procurement to sales and finance.