DiscvrAI
Customer Operations

Document + Image + Video Intelligence: The Missing Layer in Customer Operations

Text-only chatbots break down the moment a case depends on a photo, a scanned document, or a video, not just a ticket description. Real operational efficiency requires extracting and cross-referencing evidence across every format before a case is routed.

Shubham Srivastava · 21 July 2026 · 6 min read

A customer raises a damage claim. The ticket description says: "Item arrived broken, see attached." Attached are four photographs, a scanned delivery note with a handwritten annotation, and a nine-second video of the packaging.

Your text-based automation reads eight words and routes the case to a human. Everything that actually determines the outcome sits in the attachments, and the system cannot see any of it. This is the state of most customer operations automation, and it explains a great deal about why deflection rates plateau far below what the business case promised.

The evidence is not in the ticket text

In evidence-heavy operations, insurance claims, card disputes, warranty, logistics damage, field service, trade finance, the ticket text is a pointer, not the case. The determinative content is almost always in a format text automation cannot process:

  • Scanned and photographed documents, invoices, delivery notes, policy schedules, medical reports; often skewed, partially handwritten, sometimes multi-page in a single image.
  • Photographs of physical condition, damage, packaging, meter readings, vehicle state, site conditions.
  • Video, walkthroughs, unboxing, incident footage, ATM surveillance clips.
  • Structured records living elsewhere, the transaction, the shipment record, the policy master, the prior claim history.

Handling a case means reading all of these and asking whether they agree with each other. That cross-referencing step is the actual work, and it's the step nobody has automated.

Extraction is table stakes. Reconciliation is the value.

Plenty of tools will OCR a document or classify an image. That gets you fields on a screen. It does not get you a decision, because the question is never what this document says, it's whether what this document says matches everything else you know.

The interesting questions in an evidence-heavy case are all comparisons:

  1. 1Does the invoice value on the scanned document match the value in the shipment record?
  2. 2Does the damage visible in the photographs correspond to the damage type described in the claim text?
  3. 3Does the timestamp on the video fall inside the window the transaction record supports?
  4. 4Does the handwritten annotation on the delivery note contradict the customer's account of when it arrived?
  5. 5Has this claimant filed a structurally similar claim before?
A system that extracts perfectly from every format and never compares them against each other has automated the typing, not the thinking.

What a multi-format intelligence layer does

Sitting between intake and the agent's queue, its job is to reach the routing decision with the case already understood:

  • Read every attachment regardless of format, and normalise what it finds into the same structured shape as the system-of-record data.
  • Cross-reference the extracted facts against the transaction, policy or shipment record, and against prior case history.
  • Flag contradictions explicitly rather than silently picking one source, a mismatch is the single most valuable signal in the case, and it's exactly what current automation loses.
  • Attach a confidence per extracted field, so the agent knows the invoice number was clean and the handwritten date was not.
  • Route on what the evidence actually shows: complete and consistent cases to fast-track, contradictions to specialist review, missing evidence to an automated request to the customer before it ever hits a queue.

Human judgement stays where it is

None of this replaces the assessor. The case still reaches a person, and that person still decides. What changes is that they open a case where the evidence has been read, reconciled, contradictions surfaced and gaps chased, instead of one where they'll spend twenty minutes downloading attachments and squinting at a photograph of a delivery note.

The chatbot was never the constraint in customer operations. The constraint is that the hardest evidence to read is the evidence that decides the case, and almost nothing in the current automation stack can read it.

Originally published on LinkedIn.

Start with one outcome. Scale from there.

Most engagements begin as a single product on a single workflow, with a measurable result inside 8–12 weeks.