September 16, 2026

The Hidden Majority: Rethinking How Enterprises Handle Unstructured Content in the Age of AI

Every enterprise has a category of information that resists its own systems. It is not the data in the ERP or the CRM (that data was structured from the moment it was created, and the tools built to manage it work as intended). The category that resists is everything else. This is the contract, the claim file, the inspection report, the email thread with three attachments, the scanned intake form. This is unstructured content, and by most estimates it accounts for the large majority of what a typical enterprise holds.

For twenty years, the response to this has been two separate technology investments, purchased separately, run by separate teams, and rarely designed to work as one system.

Capture: the discipline of extracting data from documents

The first investment is an enterprise capture solution, or today called “intelligent document processing” or IDP. Platforms such as Tungsten Automation (formerly Kofax), ABBYY, Datacap, and OpenText Captiva ingest documents, classify them by type, and extract specific fields. This could be an account number, a total due, or a policy identifier. These are captured using OCR and template-based recognition. This is a mature, well-understood discipline, and it performs reliably against the kind of document it was designed for: high-volume, standardized, machine-printed forms.

It performs considerably less well the moment a document departs from the template. Examples would be handwriting, a form photographed, a document that varies in layout from one submission to the next, free text, and the like.  It is routinely a fifth or more of total volume, and it consumes a disproportionate share of manual review effort because every document that fails automated extraction becomes a person's problem instead of a system's problem.

Storage: the discipline of governing what you've captured

The second investment is an enterprise content management or document management system. Platforms such as IBM FileNet, OpenText, Box, and Microsoft SharePoint give the organization a system of record that includes retention rules, access controls, audit trails, and legal holds. This is a governance function, and it is essential. That is to say regulated industries do not have the option of storing content without it.

What these platforms were not built to do is understand the content they hold. Retrieval depends on the metadata applied at the point of capture. When that metadata is thin, which it often is (because tagging is expensive and rarely prioritized), the content is compliant and searchable in theory, and difficult to actually find in practice. This is the quiet cost most organizations underestimate: content that technically exists in the system of record but functions, for practical purposes, as if it does not.

Where the ground has shifted

The capability that has changed over the past three or so years is not a new module inside IDP or ECM software. It is a change in what it means to process a document at all.

Template-based systems are instructed where to look. Vision-capable AI models read the document and interpret it in a way that’s closer to the way a trained analyst would, and considerably closer than a rules-based zonal recognition engine ever could. That distinction matters most in exactly the place legacy capture has always struggled, which is the unstructured, inconsistent, and non-templated content. The improvement is not incremental accuracy on the documents that already worked. It is a materially different result on the documents that never did.

The second shift sits on the retrieval side. Semantic search and retrieval-augmented generation allow a system to answer a question posed in plain language by reasoning across the content itself, rather than depending entirely on metadata tagged at intake. This changes the economics of two decades of underinvestment in tagging discipline. Content that was captured with minimal metadata is no longer functionally lost because the system can now interpret what is inside the document, not just what was labeled about it.

Major platform vendors have begun building this directly into the environments organizations already operate. IBM's introduction of AI-driven classification and semantic search on top of existing FileNet repositories is one clear signal of the direction.  Comparable moves are underway elsewhere in the ECM market. The practical implication is that this capability is increasingly arriving as an extension of infrastructure organizations have already paid for, not as a reason to replace it.

Four applications worth prioritizing

Set against real deployments, four uses of AI in this space consistently deliver a return without requiring the organization to replace what it already runs:

  1. Extraction on non-standard documents. Handwriting, photographed submissions, scanned images, and variable-layout forms. These are the categories that have always driven exception-queue volume in template-based capture.
  2. Classification without a template library. New document types and new intake channels can be handled without engineering a new template for each one.
  3. Semantic search across existing repositories. Answering questions against content already held in FileNet, Box, SharePoint, or OpenText, without a full re-indexing effort.
  4. Synthesis for case and file review. Producing a working summary of a long file such as a claim, a loan package, or a patient record,  so a reviewer starts from an answer rather than a stack of documents.

What this does not change

AI does not remove the need for retention policy, access governance, or audit control. That discipline remains the job of the ECM layer, and none of it becomes optional because extraction got better. Nor does this shift argue for wholesale platform replacement. In most environments, the existing IDP and ECM investment performs well against the bulk of volume. The highest-return move is rarely a replatforming exercise. It’s more building or integrating a targeted layer, applied precisely to the content the current system was never built to handle, that feeds its output back into the pipeline already in place.

The distinction between organizations that get value from this shift and those that don't is not which vendor they run. It is whether they can name, specifically, which documents and which intake channels are driving their manual exception volume and whether they're willing to apply AI there first, rather than as a platform-wide initiative in search of a use case.

DAS advises regulated organizations on where AI creates the most value inside the IDP and ECM environments they already operate. Contact us at info@daspartner.com.

Key takeaways

  • Unstructured content  (documents, correspondence, images, free text) makes up 80 to 90 percent of enterprise data, yet most organizations have built their automation strategy around the other 10 to 20 percent.
  • The technology stack built to manage this content, intelligent document processing (IDP) for capture and enterprise content management (ECM) for storage and governance, was designed for predictability. Much of today's content volume is not predictable.
  • AI does not replace this stack. It closes the gap the stack was never built to close: content that has no fixed template and metadata that was never rich enough to make the content findable.
  • The organizations pulling ahead are not the ones replatforming. They are the ones identifying precisely where their existing systems break down and applying AI narrowly, against that specific gap.

Get in touch with DAS