Thomson Reuters Shows Trusted AI Starts With Data

Originally Published at Forbes on 7/6/26

AI, data analysis. Business people use AI to analyze financial related data. big data Complex performance measurement With modern innovative technology. getty

For the last two years, the enterprise AI conversation has been dominated by speed. How fast can employees draft, summarize, code, search, or analyze with generative AI? How many pilots can a company launch? How quickly can an organization say it has “adopted AI”?

But speed is a dangerous metric when the work itself depends on judgment, evidence, and trust. AI-fabricated content is not a theoretical risk. A public database tracking court and tribunal decisions involving AI hallucinations has identified more than 1,600 cases worldwide, with the count rising from roughly 700 at the start of 2026 to more than 1,200 by early April. The consequences have been severe: lawyers sanctioned, judicial decisions scrutinized, rulings overturned, and public officials forced to answer for AI-generated inaccuracies. In law, tax, compliance, and other professional services, the issue is not simply whether AI can produce an answer faster. It is whether that answer can be trusted, verified, defended, and improved over time.

That is why the first wave of enterprise AI adoption is beginning to look shallow. A lawyer producing a memo faster does not help much if the underlying answer cannot be trusted. A tax professional moving more quickly does not matter if the AI misses context buried in a complex document structure. An enterprise deploying agents at scale will not get far if those agents keep making the same mistakes and no one has built a system for improving them.

The hard part is no longer getting access to AI. It is making AI good enough to rely on when mistakes have real consequences. That is why case studies where enterprises are not just adopting AI, but improving it over time, are worth paying attention to. Thomson Reuters’ work with Invisible Technologies highlights the critical infrastructure required to make AI reliable in professional workflows, while also showing how external AI data and training expertise can complement the deep domain expertise of Thomson Reuters legal, tax, accounting, and compliance experts. Invisible helps Thomson Reuters move quickly on specific data workflows and adapt to new task types, while Thomson Reuters subject-matter experts continue to provide the domain-specific judgment required for professional-grade AI.

What makes the example useful is that Thomson Reuters is not approaching AI as a generic productivity tool. It is building for professions where the product depends on accuracy, evidence, and professional judgment: legal, tax, accounting, compliance, and news. Its starting point is trusted content, deep domain expertise, and established workflows in fields where “mostly right” is not good enough.

As Joel Hron, Chief Technology Officer at Thomson Reuters, put it, the company is building AI for customers who need outputs they can trust, verify, and defend in high-stakes professional settings. Thomson Reuters calls this as Fiduciary-Grade AI™: systems designed for professional environments where accuracy, transparency, and accountability are essential. For these customers, AI has to do more than generate plausible answers. It has to understand the underlying structure and meaning of complex professional documents.

A lot of companies are still treating AI as a layer they can place on top of existing workflows. Thomson Reuters is going deeper. Its CoCounsel capabilities are being applied across legal and tax work as Thomson Reuters evolves its products for the agentic era. Hron described the broader leadership challenge: Thomson Reuters has had to determine where AI should improve existing products and workflows, and where it should enable entirely new ways of working across legal, tax, accounting, compliance, and news.

That is a very different leadership problem from “How do we add a chatbot or co-pilot?” It is the problem every large company will soon face: which products, workflows, roles, and decision systems should be redesigned because AI changes the underlying economics of the work, and how can AI be made reliable enough to support those systems?

One specific example sits underneath the broader transformation. Thomson Reuters needed high-fidelity training data for a document layout segmentation model capable of interpreting visual structure and contextual meaning across diverse documents, such as invoices and forms. Traditional OCR (Optical Character Recognition) can extract text from a document, but it often misses the structure that gives that text meaning. In complex documents, tables, form fields, nested questions, and multi-part questionnaire responses are not just formatting details; they define how information should be interpreted. The model needed to understand those relationships, not simply transcribe the words.

Invisible built a flexible labeling system that could capture every element of a document, from tables and forms to nested fields, mapped with precise, pixel-level accuracy. Invisible was well positioned to do this, after working on similar content extraction training at major labs. It also supported a workflow for expert annotation and review within Invisible’s platform, helping ensure datasets were consistent and ready for use in Thomson Reuters model development process. After the initial work was delivered, Thomson Reuters expanded the engagement to support additional data needs. The leadership lesson is simple: in high-stakes AI, the model is only as useful as the context it can reliably understand.

Many enterprises have deployed agents or AI workflows, but few have built the feedback systems required to make them better. They make the same mistakes repeatedly because no one has designed a mechanism for feedback, post-training, model improvement, or disciplined use-case refinement. Thomson Reuters, by contrast, is investing in the less glamorous layer that separates demos from durable enterprise capability: improving models with better data, expert review, and specific feedback loops.

One example of that discipline is CoCoBench, Thomson Reuters internal legal AI evaluation standard. CoCoBench is designed to test whether an AI system can complete real legal tasks to the standard required for Fiduciary-Grade AI™, not just answer isolated prompts. It is built on hundreds of attorney-authored tasks across research, drafting, review, and revision, each paired with a gold-standard response written and reviewed by practicing attorneys. More than 100 legal subject-matter experts and Thomson Reuters Labs researchers contributed to its development, representing more than 15,000 hours of work, and Thomson Reuters uses a fixed core dataset to track performance over time.

This is where many companies are getting stuck. They are throwing large language models at entire workflows when only a small part of the workflow actually requires an LLM. They are using token-heavy systems where deterministic software would be cheaper and more reliable. They are launching broad AI access without deciding which use cases deserve deep investment, training, and governance.

Invisible’s view, based on its enterprise work, is that more mature organizations are becoming more selective. Rather than deploying 100 mediocre AI use cases, they identify the handful that truly move the business and make those excellent. That requires asking a detailed set of questions: Where does AI actually create enterprise value? Where does the work require a probabilistic model? Where would deterministic software be better? Where do human experts need to remain in the loop? Where is the cost of error too high to tolerate generic AI output?

Thomson Reuters’ answer starts with trust. The company has more than 2,600 subject-matter experts who help shape, evaluate, and improve its content, products, and AI systems. That human layer is not a sign that the AI is weak. It is what allows AI to become useful in professional contexts where hallucination, missing context, or a flawed citation can create real consequences. Partnerships like the one with Invisible add another layer of capability: flexible, high-quality feedback and annotation workflows that help Thomson Reuters move quickly across new task types, while preserving the deep legal, tax, accounting, and compliance expertise that differentiates its products. That is the next misconception leaders need to drop: human review is not a temporary bridge until the technology gets better. In consequential work, expert feedback is part of the architecture.

The best enterprise AI systems will route, augment, test, and improve human judgment. Humans will provide the domain expertise, edge-case judgment, validation, and accountability that models cannot carry alone. The AI will accelerate the work, surface patterns, structure complexity, and reduce manual burden. But the system around both will determine whether the output is merely faster or actually better.

That is also why the productivity conversation needs to mature. Time savings are useful, but time saved is not the same thing as value created. If AI helps individuals work faster, but the organization does not redesign workflows, decision rights, collaboration patterns, or business metrics, the benefit can remain trapped at the individual level. The organizational ROI gap often appears when AI is deployed as a personal productivity tool rather than as an enterprise operating model shift.

Thomson Reuters understands that. The company is not merely giving professionals AI assistants. It is asking what professional-grade AI requires underneath: authoritative content, advanced document understanding, rigorous evaluation, expert review, trusted outputs, and products designed around agentic workflows. There is another lesson here for CEOs, CIOs, and product leaders: the AI-first enterprise will have to become much more disciplined about what it builds.

The first era of enterprise AI rewarded experimentation. The next era will reward judgment. Leaders will need to know when to use a general-purpose model, when to fine-tune, when to post-train, when to build deterministic software, when to keep humans in the loop, and when to retire products or workflows that AI has made obsolete. That is a harder mandate than launching pilots. It requires product strategy, data strategy, governance, talent strategy, and operating model design to come together.

It also changes what companies need from their people. Hron noted that Thomson Reuters technology leadership is thinking differently about hiring engineers. Domain expertise still matters, but flexibility, cross-functional range, and the ability to work across boundaries are becoming more important as AI changes how work gets done.

That may be one of the most overlooked effects of AI. As tools become more powerful, old swim lanes start to blur. Engineers, product leaders, marketers, legal experts, and operators can do more outside their traditional lanes. Done poorly, that creates confusion and duplicated effort. Done well, it can break down silos and accelerate enterprise learning. But again, the difference is design.

The work Thomson Reuters is doing with Invisible offers a useful reminder that the success of AI transformation depends on the invisible infrastructure behind the prompt: the document schemas, pixel-level annotation, expert review workflows, and training data pipelines that make AI dependable enough for real work.

That is where enterprise AI is headed. And for leaders still asking how to “adopt AI,” the better question is becoming: Do we have the data, review systems, and operating discipline to make AI trustworthy in the workflows where accuracy matters most?