AI工具Score A (74)

Harvey + Legora on OpenAI's GPT-6 Astra - Artificial Lawyer

2 小时前2 viewsSource: artificiallawyer.com
OpenAI has released GPT-6 Astra, its new flagship model and their most capable LLM yet for end-to-end tasks. Both Harvey and Legora have provided some insights into how it benefits legal AI, (see below). GPT‑6 Astra started its roll out on September 3rd on a limited basis and will eventually be more widely available. But, the question is: with so many models arriving now, is this a big deal? The short answer is: yes, because the major model makers are showing that they can – albeit incrementally – keep improving AI performance. And that matters especially for lawyers, as there is a hard to define threshold under which legal AI is of limited standalone use, but conversely where once it hits a certain level of dependability it becomes immensely powerful. For example, OpenAI provided a range of benchmarks showing many improvements, one of which was ‘Agents’ Last Exam’ – a benchmark created at UC Berkeley RDI to answer the question: ‘ Can an AI agent perform economically valuable professional work from beginning to end ?’ It covers 55 professional ‘subdomains’ and includes some aspects of legal work, among many others. In this case (see graph) it reached 59.3%, while GPT-5.6 Sol reached 53.6%. OpenAI data and graph, 2026. Now, you may then say: ‘So what? That’s not good enough to fully take over a whole stream of legal work.’ And you would be correct. But, it’s getting better – quickly. If one considers that when genAI was first applied to legal tasks the idea of an LLM handling a whole stream of complex work – and doing it well – seemed like a pipe dream; so we have come a long way. There is no indication that things will slow down, especially given the huge competition now between both the main model makers, and also the open-source developers. OpenAI added: ‘ GPT‑6 Astra pairs advances in computer use with targeted training for professional environments , to help tackle complex work tasks. It combines the intelligence required for complex problems with the ability to carry out multistep workflows and produce polished documents… ‘Astra is our most aligned model, with substantial improvements in understanding user intent and model behaviour, you can delegate tasks with greater confidence in Astra’s judgment .’ And all of that matters if we are going to take a truly agentic approach to legal work, i.e. not just a prompt here and there, but asking the AI tool to perform a range of connected tasks. OpenAI also provided a couple of examples of formatting specifically for legal documents, but those are not that exciting, even if formatting remains a pain for lawyers all over the world. — Advertisement — — Harvey and Legora’s View Harvey’s Niko Grupen , Head of Applied Research at the fast-growing legal AI platform, had this to say: ‘Astra is a significant quality improvement over GPT‑5.6 Sol across complex legal tasks. ‘In our early testing, Astra stood out by approaching legal work the way a discerning lawyer does : it distinguishes documents from established records, surfaces unsupported assumptions, and converts gaps into concrete drafting positions.’ AL looks forward to seeing how it performs on Harvey’s own benchmarks. For Legora, OpenAI published a post with legal engineer Percevale Perks , which involved a trial of GPT-6 Astra on what they described as ‘one of the more tedious workflows … financial-statement tie-out: checking every figure in draft accounts against trial balances, a consolidation schedule, and the previous year’s accounts until each item agrees.’ They aimed the new model at it and this is what they found: ‘ Processing complex financial context at scale – ‘Using GPT‑6 Astra, Legora’s Agent completed the tie-out across 41 documents in a single run. The Agent did the work within minutes: checking every balance against its supporting schedule, surfacing breaks in the amounts, and recording each check. The result gives the legal professional a granular record of every line item and figure to review. ‘I think what changed before and after is the processing power, the ability to ingest such a large number of documents, digest really complex information, and get all of those different line items and figures. ‘The Agent handles the exhaustive comparison, while the expert remains responsible for the judgment call on each result. That approach of keeping a human in the loop is central to how Legora is extending its platform beyond legal work into audit, tax, compliance, and risk. ‘ GPT‑6 Astra found all four errors Legora had planted in the accounts, including a £500,000 gap hidden in the revenue note. It checked every balance against its supporting schedule and recorded each check. And it retained every check the previous model got right as well as completed around 50 more. ‘The result is a more complete and faster first pass and a clearer record to review, while the final decision stays with the legal expert.’ Key here is the time saved, the complexity of the task, and the sense of confidence it gave in the results. Although, in this case, this test was more about numbers than complex legal text. Legora also evaluated GPT‑6 Astra with its Legora Benchmark for Agentic Reasoning (BAR), which measures performance on end-to-end legal tasks drawn from real-world use cases. Legora found that: ‘GPT‑6 Astra improved performance by nearly 40% over the previous model on this financial-statement workflow. Across all tasks in the BAR, the improvement averaged about 3%.’ Conclusion Is this an Earth-shattering leap? No. Is this proof that models can meaningfully advance every few months? Yes. Why does this matter? Because legal work is highly dependent on accuracy. If an agent can take on a complex legal workflow and with such high quality results that lawyers can truly rely on them with just the lightest of quality checks, then that truly changes the game. Are we there yet? No. As seen above, on Agents’ Last Exam, the new model got 59.3%. That is not a ‘fire and forget’ level of performance. But, the rate of improvement is notable. In a few days no doubt another company will release another model that takes things on a little more. And then another one will. And so on. Slowly the climb up the performance curve will advance. Does this have to reach 99.9%? No. One could argue it’s already very useful as it is, as long as there is some experienced lawyer supervision. The main question is at what point does performance on complex agentic tasks get so good that the lawyer is just ‘giving things a quick once-over to be sure’? And at that point how much does this then change the legal business model more widely? More about Astra here . — Come and join us in London this November 4th + 5th at Legal Innovators ! The conference at the intersection of legal AI and the business of law. Day One: law firms, Day Two: inhouse. Legal Innovators UK – London, Nov 4 and 5 — Share this: Tweet Click to email a link to a friend (Opens in new window) Email Discover more from Artificial Lawyer Subscribe to get the latest posts sent to your email. Type your email… Subscribe

Read the full original article:

artificiallawyer.com
#OpenAI