Cross-Document Correlation: Reconciling Fuel Dockets Against Weighbridge Tickets Against Meter Readings
The reasonable assurance walk-through samples a single emission record and asks whether the source docket reconciles with other records in the same period. Here is how automated cross-document correlation surfaces mismatches before the auditor does.
An auditor pulls a fuel record from Q2. It shows a 4,200 litre bulk diesel delivery to a Pilbara site on 14 April. Reasonable question: where is the weighbridge ticket that confirms the tanker arrived? Where is the site log that shows the tank went from 8,000 litres to 12,200 litres that afternoon? Where is the fuel card statement that shows nothing was drawn from that supplier under a separate account for the same load?
Single-document extraction cannot answer any of those questions. It pulls the number off one PDF and stops.
That is the gap Cross-Document Correlation is built to close. Not extraction. Reconciliation between extractions.
Why single-document extraction is not enough for reasonable assurance
Year 2 assurance under AASB S2 moves from limited to reasonable for the same Group 1 entities that reported in FY25. The evidence bar is different. Limited assurance samples lightly and looks for obvious contradictions. Reasonable assurance samples deeply and looks for reconciliation.
The auditor's question changes from "does this number look right on its own" to "does this number agree with the other numbers about the same physical event".
A fuel purchase is not one number. It is a delivery quantity on a docket, a weight change at the weighbridge, a tank level change on the site log, a payment on the bank statement, and eventually a line on the supplier's monthly statement. Reasonable assurance expects those to line up. If they do not, the auditor will not ignore it. They will ask what happened, and if the sustainability team cannot answer, the finding goes into the management letter.
Most carbon platforms extract each document in isolation. The number lands in the ledger. Nothing checks whether it agrees with the other numbers about the same physical event. That is the reconciliation gap.
The four correlation patterns
The correlation engine runs four checks across every organisation's document corpus. Each one maps to a real reconciliation the auditor expects the reporter to have already done.
1. Fuel docket against weighbridge ticket (bulk deliveries). A bulk diesel delivery to a mine, quarry or construction site typically generates two documents: the supplier's delivery docket and the site's weighbridge ticket. Both describe the same physical event, from opposite sides of the fence. If the docket says 4,200 litres and the weighbridge ticket implies 3,850 litres, the reporter has a choice, pick one, reconcile the two, or flag it. The correlation check catches when both documents land in the ledger against the same material and the quantities disagree, or when only one of them has landed.
2. Fuel card statement against fuel card individual transactions. Ampol, BP and Shell issue monthly statements that summarise every card transaction. They also issue individual transaction records. When both flow into the ledger, the risk is double counting, the monthly total lands once, then the individual transactions land again. The correlation check flags any period where the same material and quantity appear on two documents against the same supplier.
3. Utility bill against smart meter reading against sub-meter allocation. A shopping centre landlord receives a whole-of-site electricity bill. The tenant receives a sub-meter allocation for their portion. The building management system exports smart meter data for the same period. All three should agree, the tenant's sub-meter total plus other allocations should reconcile with the whole-site bill. The correlation check surfaces the case where a tenant's allocation and the whole-site bill both land against the same facility, or where the smart meter export and the retailer bill disagree by more than a rounding error.
4. Refrigerant service docket against refrigerant management log. A refrigeration technician's service docket shows 2.4 kg of R-404A charged into a chiller during a June service. The site's own refrigerant management log, the annual leak register, should show the same 2.4 kg addition. Under NGER's fugitive emissions rules, these are the two evidence points an auditor cross-checks. The correlation engine catches the case where one landed and the other did not, or where the same charge event appears twice from two different data feeds.
Those are the four practical patterns. Under the hood, they resolve into four generic detection algorithms: duplicate invoice, overlapping period, same-supplier quantity mismatch, and identical-quantity double-count within a month. The algorithms do not know they are looking at fuel dockets or refrigerant logs. They know they are looking at extractions that share a material, a supplier, or a period and that should not both exist independently.
How the correlation runs
The scan is background work, not a modal blocker. When a document finishes extraction and lands in the ledger, it does not need to wait for reconciliation before the emission record commits. The correlation engine runs on a rolling window, by default the last 90 days of extractions for that organisation.
Inside that window it fetches every extraction with a material, a quantity, and a date. It orders them, indexes them by material and by supplier where a supplier is present, and then runs the four checks. The output is not a passing score. The output is a list of findings, each one pointing at exactly two documents and the specific fields that clash.
A finding is scoped to the organisation. It never crosses tenant boundaries, one client's fuel docket cannot correlate against another client's weighbridge ticket, even if the same haulage contractor delivered both. That is a hard architectural rule.
The engine also remembers what it has already found. If two documents were flagged as a duplicate invoice last week and the reviewer resolved the finding, the same pair does not re-appear the next scan under the same correlation type. That deduplication is stored in the database and checked before any new findings write.
What mismatches surface
Four correlation types show up in the review queue.
Duplicate invoice. The same material, quantity and date appear across two different documents. Severity: high. This is the pattern that catches the fuel card statement double-counted against its individual transactions, or the same supplier invoice ingested twice through email and OneDrive.
Overlapping period. The same material extracted from two different documents within seven days of each other. Severity: medium. This is the pattern that catches a delivery docket and a weighbridge ticket for the same load. It is also the pattern that catches a genuine, legitimate case where one facility took two deliveries in the same week, which is why the finding goes to review, not to auto-resolve.
Same-supplier mismatch. The same supplier and material appear on multiple documents but the quantities differ by more than 10x. Severity: medium. This is the pattern that catches unit errors. Litres logged as kilolitres. Cubic metres logged as litres. Kilowatt-hours logged as megawatt-hours. The 10x threshold is deliberate. A real business does not suddenly buy ten times as much diesel from the same supplier in the same month without an explanation.
Quantity inconsistency. An identical quantity for the same material appears in the same calendar month from two different documents. Severity: low. This is the pattern that catches a template refuel amount (a driver who always writes "50L" on the docket regardless of the actual pump reading), or a suspicious round number appearing on both a docket and a card statement.
Each finding lands in the review queue with a description, a severity, a similarity score where relevant, and the two document IDs. The reviewer sees both documents side by side and the fields that clashed.
The reviewer investigation workflow
A finding is not a defect. It is a question that needs an answer written down. That is the framing that survives assurance.
The reviewer opens the queue, sees a finding, and has four resolution paths.
Reconcile. Both documents describe the same real event, and one of them is authoritative. The reviewer marks which one, retires the duplicate from the ledger, and writes one sentence explaining why. That sentence becomes part of the audit trail.
Confirm distinct. The two documents describe two real events that happen to look similar, two legitimate deliveries in the same week, or a coincidence of round numbers. The reviewer marks distinct and writes why the coincidence is genuine. The finding closes, and the same pair does not re-appear.
Escalate. The mismatch is real and the reviewer cannot resolve it alone. It goes to a manager, an operations contact at the site, or a supplier contact. The correlation record links to the escalation.
Correct. Both documents were right; the extraction was wrong. The reviewer fixes the underlying emission record, adds a restatement note, and the correlation closes.
None of these four paths involve deleting the finding. Every finding either resolves to reconciled, distinct, escalated or corrected, and every resolution is timestamped, attributed to a user, and preserved in the audit log for the seven-year retention window.
Integration with the Anomaly Detection queue
Correlation mismatches do not sit in a separate silo. They flow into the same review surface as agentic anomaly findings. One queue, one investigation workflow, one resolution log.
That matters because a reviewer working through Q2 preparation does not want to check three separate tools for three separate lists of things to investigate. They want one queue, sorted by severity, that says: here is everything the platform is not confident about.
Anomaly detection catches statistical outliers, a fuel record five times its historical average, a refrigerant charge that violates a rule of physics. Cross-document correlation catches structural conflicts, two documents that describe the same event and disagree. Together they cover the two failure modes an auditor probes for during walk-through. Neither one is sufficient alone.
How this passes ASSA 5010 walk-through testing
The Auditing and Assurance Standards Board's ASSA 5010 sets out what a reasonable assurance engagement over a sustainability report actually involves. The critical section for our purposes is walk-through testing. The auditor picks a small sample of emission records, usually 25 to 40 across scopes, and traces each one from the ledger back to source.
The walk-through does not stop at the source document. For selected records, it extends into reconciliation. The auditor asks: was this delivery cross-checked against the weighbridge? Was this fuel card statement cross-checked against transaction detail? Was this refrigerant charge cross-checked against the leak log?
If the answer is yes, and the reconciliation is preserved in the audit trail with a timestamp, a reviewer identity and a written resolution note, the walk-through moves on. If the answer is no, the auditor either extends the sample or issues a control deficiency. Extended samples cost the reporter, every additional record traced is billable audit time. Control deficiencies cost more.
Automated correlation flips the economics of that walk-through. The reconciliations are already done. They are already documented. The auditor sees a resolution note next to every flagged pair. Sample extension becomes unnecessary in the vast majority of cases.
Combined with a durable emissions data lineage trail, this is the difference between a walk-through that takes an afternoon and one that spills into a second week.
What the platform does NOT do
Two honest disclaimers.
The engine does not perform blockchain-style tamper detection. Some vendors market cryptographic hash chains and immutable ledgers as the answer to audit trail integrity. We hold content fingerprints on every source document and every emission record, and we log every write, but we do not stake the audit trail on a distributed ledger. The reviewer identity, timestamp and content hash are enforced at the database layer. That is what auditors we have talked to actually want. It is not what the marketing category promises.
The engine does not automatically resolve mismatches without human confirmation. Every finding requires a human resolution. There is no auto-merge, no silent deduplication, no confidence threshold above which the platform decides for the reviewer. If a finding cannot be closed with a written resolution note, it stays open. That constraint is deliberate. An auditor accepts "the reviewer determined X and here is why" as evidence. They do not accept "the system determined X" as evidence of anything.
What to try
The correlation engine runs on the same document corpus per-supplier extraction templates build up. The richer the extraction, the sharper the reconciliation. Load a real quarter of your operations, fuel dockets, weighbridge tickets, utility bills, refrigerant service reports, and let the scan run overnight. The findings queue in the morning is a preview of what your Year 2 assurance auditor will ask about.
Per-project pricing plus $100 per month per organisation, with the Auditor Workspace, Evidence Pack export and correlation engine included at every tier. Email hello@carbonly.ai and mention which reconciliations you want tested first.
FAQ
How is cross-document correlation different from anomaly detection?
Anomaly detection catches a single record that looks statistically wrong on its own, a fuel consumption reading five times last month's, or a refrigerant charge that exceeds the equipment's total capacity. Correlation catches two records that look reasonable individually but conflict when compared. Different failure modes, same investigation queue.
Does the engine work across organisations, so we can catch a supplier double-billing across our subsidiaries?
Correlation is scoped to a single organisation by design. That is a data isolation boundary, no client's documents are ever compared against another client's. However, JV consolidation and multi-entity organisations run correlation across all consolidated entities within their own boundary, so a group with ten operating sites gets group-wide reconciliation without the platform ever crossing into another customer's data.
What happens if I do not resolve a finding? Does it stay open forever?
Yes. There is no auto-close and no time-based expiry. An unresolved correlation finding is a documented open item that appears in the auditor's review pack. If it is genuinely a false positive, mark it "confirm distinct" with a one-line explanation. That takes five seconds and preserves the audit trail. Leaving findings unresolved is the worst option, it looks like ignorance rather than judgement.
How far back does the scan look by default?
Ninety days of extractions, on a rolling window. Longer windows can be run for specific investigations or for the pre-assurance sweep at reporting period end. See the three questions your assurance provider will ask that a spreadsheet cannot answer, one of them is exactly this: what is your reconciliation window across source documents.
Can the correlation engine catch a supplier issuing two invoices for the same delivery under different reference numbers?
Yes, provided both invoices carry the same material, quantity and delivery date. That is the duplicate-invoice check and it triggers regardless of the invoice numbers being different. If the supplier issues two invoices for legitimately different deliveries that happen to have the same quantity in the same month, the reviewer closes it as distinct and the pair does not re-flag.
For teams working through remediation, the correlation queue often feeds directly into compliance gap fixes and agent-suggested remediations. One system finds the mismatch; the other proposes how to close it.
Related reading
- Agentic Anomaly Detection: Five Rule Types That Catch Errors Auditors Otherwise Find
- Three Questions Your Assurance Provider Will Ask That a Spreadsheet Cannot Answer
- Emissions Data Lineage: Traceability From Source Document to Ledger
- Per-Supplier Extraction Templates: Boral, Ampol, AGL and the Rest
- Compliance Gap Fixes: Agent-Suggested Remediations for Carbon Accounting