Agentic Anomaly Detection in Carbon Accounting: How Five Named Rule Types Catch the Errors Auditors Would Otherwise Find
Traditional anomaly detection in emission ledgers is a rule table maintained by a data engineer. Here is how a five-rule agent runtime, priority scoring, grouping and a trust loop catches the data quality issues an ASRS auditor would otherwise find during walk-through.
Year 1 assurance under AASB S2 was limited assurance. Year 2 for the same Group 1 entity is reasonable assurance. The sample size the auditor pulls goes up. The tolerance for unexplained variance goes down. The bill for internal walk-through preparation is the same 40 to 60 hours per site, except now it lands twice.
Most sustainability teams are still finding the errors during that walk-through. Not before. That is the wrong end of the pipeline to be looking.
Anomaly detection, done properly, is the layer that catches the numbers your auditor would otherwise circle in red. It runs on every write. It ranks what it finds. It bundles similar issues so a reviewer sees one investigation instead of five hundred. And when the same resolution pattern shows up enough times, it learns.
This is a walk-through of how we built that layer inside Carbonly, and specifically how five named rule types map to the errors that turn up in real Australian emission ledgers.
Why anomaly detection matters more under reasonable assurance
Under ASSA 5010 limited assurance, the auditor is asking whether anything has come to their attention that suggests the disclosures are materially misstated. Under reasonable assurance in year 2 and beyond, the auditor is forming a positive opinion. That is a different exercise. It means substantive testing across a wider sample of your emission records.
If the sample includes a duplicate diesel invoice, a mid-period consumption spike that nobody investigated, or a Scope 2 factor from a superseded NGA edition, that ends up on the management letter. If the same class of issue shows up more than once, it can drive a qualified opinion.
The unit of prevention is a first-class anomaly object. Detected at write time. Ranked by likely impact. Investigated with evidence attached. Resolved through a typed workflow the auditor can walk. That is what the Anomaly Detection module in Carbonly gives you.
The five rule types
A rule table in a spreadsheet, maintained by whoever last had capacity, is not a data quality strategy. The five rule types below are the ones our detection engine runs against every emission ledger write.
1. Statistical outliers
For every material and every facility with at least ten historical records, the engine computes a mean and a standard deviation of the quantity value. A record with a z-score greater than 3 gets flagged medium. Greater than 5 is high. Greater than 8 is critical.
We also run rolling-median deviation for materials where the underlying distribution is skewed (fuel consumption almost always is, because zero-litre months exist and pull the mean around). The rolling median gives a more honest baseline than a pure z-score when the historical series has legitimate gaps.
The check runs on the extracted quantity, not the calculated tCO2e. That matters. A data-entry error that puts 50,000 in the litres field instead of 5,000 is caught before the emission factor is applied. If you only outlier-check the tCO2e result, you have already burned the mistake into your ledger.
2. Rate-of-change flags
A monthly natural gas bill that runs at 15,000 GJ for eight months and then drops to 3,000 GJ in month nine is either a meter fault, a submission error, or a real event you need to explain to the auditor. The rate-of-change rule flags mid-period movement that exceeds the historical variance envelope for that material and facility.
The math is boring. A 30-day rolling window is compared against a 90-day baseline. Deviations beyond a configurable multiple of the standard deviation trigger a flag. What makes it useful is that the flag carries the historical series with it, so the reviewer opens the anomaly and sees the pattern they need to explain in one screen.
3. Threshold flags
Some values are physically implausible. A single fuel record for 100,000 litres is almost certainly the site's storage tank volume mistyped into the delivery quantity field. A refrigerant top-up exceeding the original charge on the plant nameplate is either a leak the site did not report or a units confusion.
The threshold rule runs a set of hardcoded plausibility bounds per unit type. Kilowatt-hours above ten million per record. Litres above one hundred thousand. Tonnes above ten thousand. Cubic metres of natural gas above one million. Zero and negative quantities are flagged separately at higher severity, because they are almost never legitimate and they distort every aggregation downstream.
These are not statistical rules. They are physical-world rules. They catch the errors that the outlier check misses because the historical baseline was already broken.
4. Cross-reference flags
The activity record from a supplier invoice ought to reconcile with the procurement PO, the delivery note, and the finance ledger accrual. When it does not, that gap is the auditor's first question and your first opportunity.
The cross-reference check groups emissions by supplier and month, then flags any single-document contribution that deviates more than 50 percent from the group's per-document average. For a subset of flagged groups, the check runs a contextual pass to explain the inconsistency. Duplicate submission. Credit note. Legitimate second delivery. The reviewer sees the group and the suggested classification together.
This is the check that pays for itself the first time a supplier double-bills a period and nobody upstream in AP catches it before it lands in your emission total.
5. Factor version drift
Every year DCCEEW publishes an updated NGA Factors workbook. NGER-covered facilities have to use the correct edition for the correct reporting year. Historical records that still apply a superseded factor are a specific class of error that is easy to miss because the record itself looks fine.
The factor version drift rule compares every tenant-configured emission factor against the current NGA library. When a divergence is detected, the anomaly carries the old factor value, the new factor value, and the count of affected records. The reviewer can accept the update, override it with justification, or restate the period.
This one is quietly the highest-impact of the five. A single Victorian grid factor moving from a superseded edition to the current one restates every Scope 2 electricity record for the year.
Priority Scoring: how the queue ranks
Detection is only half the problem. If the queue is a flat list of 518 items sorted by created-at, reviewers will work the top of the list until they get tired and the actually-important anomaly at position 340 will sit unaddressed until the auditor finds it.
The Priority Scoring service produces a single 0 to 100 score per anomaly from four weighted components:
- Severity weight (35 percent). Critical, high, medium, low maps to a fixed score band.
- CO2e impact weight (30 percent). The tCO2e magnitude of the issue if it is real, scaled against the organisation's total emissions. A 500 tCO2e anomaly at a site with 800,000 tCO2e annual is different from the same 500 tCO2e at a site with 3,000 tCO2e annual.
- Confidence weight (20 percent). Some detection methods are more reliable than others. An exact duplicate match is high confidence. A statistical outlier at z-score 3.1 is lower.
- Age weight (15 percent). Recent anomalies score higher, decaying over 30 days. This keeps the freshest problems visible.
The reviewer opens the queue and sees the anomalies that most deserve attention today at the top. The trust orchestrator uses the same score to decide which items are safe to auto-resolve, which we get to shortly.
Anomaly Grouping: N similar issues, one investigation
If a supplier double-submits a fuel invoice run across 47 sites, a naive detector produces 47 anomalies. A reviewer working item by item wastes a day.
The grouping service clusters anomalies on four rules. Same rule and same project becomes one group. Same duplicate pair becomes one group. Same missing-data material across months becomes one group. Same trend direction and scope becomes one group.
Each group returns a summary card with the count of member anomalies, the total CO2e impact, the worst severity, and the most recent anomaly as the representative. The reviewer opens one group, applies one resolution, and the resolution propagates to every member.
This is where the internal walk-through hours come down. Not by hiding work, but by presenting it at the right altitude. One decision instead of 47.
Anomaly Investigation: evidence before opinion
Opening an anomaly should not drop the reviewer into a blank form. It should show them what the system already knows.
The investigation service dispatches on the rule type and produces a structured result with four fields. Evidence is the set of specific data points that support the finding: for a duplicate, the matching records with the fields that differ highlighted; for a threshold breach, the historical trend and the deviation percentage; for factor drift, the old and new factor values plus the affected record IDs.
Reasoning is the plain-English explanation of what was found. Confidence is a 0 to 1 score indicating how strongly the evidence supports the finding. Suggested actions is the list of resolutions available for this rule type.
The suggested actions are typed. For duplicates, they are mark_duplicate, exclude_from_report, or merge_records. For missing data, send_data_request or create_estimate. For threshold and outlier findings, create_incident or flag_for_review. For factor drift, apply_updated_factor.
None of the investigation logic requires a language model. It is pure data lookup and deterministic analysis. That means every run is reproducible and every finding can be walked backwards by an auditor asking "how did the system arrive at this?"
Anomaly Resolution and the trust loop
Every resolution action is auditable, reversible, scoped to the organisation, and non-destructive. The agent never hard-deletes an emission record. It soft-marks, flags, or estimates. Every modification writes to the audit log with the reasoning, the evidence used, and the reverse operation for undo.
The Trust Orchestrator sits above this and enforces the three-mode graduation model that runs across every autonomous surface in the platform. In Shadow mode the agent enriches every anomaly with investigation and priority score but takes no action. The reviewer sees a shadow report of what the agent would have done. In Copilot mode the agent enriches and presents findings with suggested actions; the reviewer approves or rejects. In Trusted mode the agent auto-resolves anomalies whose CO2e impact falls below a rule-configured threshold and whose suggested action is on the rule's allowlist.
A per-hour rate limit sits on top so a runaway agent cannot mass-modify data. Every auto-resolution is a typed audit event with a reverse instruction attached.
The trust loop is what makes the module get faster over time. When a resolution pattern recurs (say, the same supplier's duplicate invoices always resolve to mark_duplicate with the same rationale), the orchestrator can graduate that pattern from "flag every occurrence" to auto-suppress with the standing rationale attached. Every graduation is itself a typed event on the audit ledger. The auditor walks the decision the same way they walk any other resolution.
How this compresses the walk-through
The 40 to 60 hours per site the sustainability team spends preparing for the auditor's walk-through breaks down roughly into: pulling source documents, reconciling ledger to source, explaining variances, and documenting exceptions. Every hour of that is an hour the anomaly module has already spent, provided you have been running it during the period.
The auditor's first question in walk-through is nearly always some version of "what data quality issues did your system catch before you submitted?" A first-class anomaly object with a resolution workflow and a typed audit event is the answer. Not a spreadsheet. Not an email trail. A queryable record with evidence attached.
The reduction is not automation of the auditor's work. It is showing the auditor the work you did during the period, in the format their file already wants.
What the module does not do
Being honest about the limits keeps expectations calibrated.
Root cause at the operational level. If the flag is a fuel consumption spike, the module tells you the spike is real and quantifies it. It does not tell you the site had a compressor failure that week. That analysis stays with the site team. The module gives them a clean pointer to the record and the historical context. It does not walk the plant.
Greenwashing intent. Whether a claim is misleading is a policy question, not an anomaly. The module can tell you your Scope 2 factor is stale, but it cannot tell you your net-zero-by-2035 claim is unsupported by your transition plan. That is a governance decision, not a data-quality flag.
The auditor's professional scepticism. Anomaly detection surfaces the issues before the auditor arrives. It does not substitute for the auditor. Reasonable assurance still requires an independent view, and the auditor will still sample beyond what the system flagged.
The Claude and ChatGPT surface for the CFO and Head of Sustainability
Most CFOs do not open the Anomalies page. They ask a question in the tool they already have open, which increasingly is Claude Desktop or ChatGPT.
Carbonly exposes anomaly data through the same MCP integration that surfaces the rest of the ledger. A Head of Sustainability can connect Claude or ChatGPT to the workspace and ask questions like "what anomalies did we flag this quarter" or "show me any factor drift on Scope 2 electricity". The response comes back as a ranked list with priority scores, CO2e impact figures, and links back to the source records. The integration respects the same RBAC as the web app, so the AI assistant sees exactly what the user is entitled to see and nothing more.
The point is not to move the work into chat. It is to remove the excuse of "I did not know we had that issue" from the CFO's monthly review.
Consultant angle
External sustainability consultants running quarterly data quality reviews often start from a blank Excel workbook and build a data profile from scratch. That is billable time that produces nothing durable.
Consultants running client work on Carbonly start from the anomaly queue. The heavy lifting (statistical baselines, cross-document reconciliation, factor drift detection) has already run. The consultant applies their judgement to the ranked list and the grouped investigations, which is what the client is paying for. The workshop equipment is already in the room; the craftsperson brings the craft.
Related reading
- Anomaly detection in emissions data: what it catches and why it matters
- Agentic AI workflows in carbon accounting
- AI carbon accounting, self-learning and the human-in-the-loop for accuracy
- Autonomous AI agents in carbon accounting
- Data quality in carbon accounting: what auditors actually check
- NGER compliance automation
- ASRS assurance requirements: what the auditor will ask for
- The AI pipeline behind utility bill extraction
FAQ
What kinds of anomalies does Carbonly detect? Five rule types run on every emission ledger write: statistical outliers (z-score and rolling median), rate-of-change flags on mid-period consumption movement, threshold flags for physically implausible values, cross-reference flags where activity records do not reconcile across documents, and factor version drift where a record still applies a superseded NGA edition. Each finding becomes a first-class anomaly object with evidence, priority score and a typed resolution workflow.
Does the AI ever suppress a real anomaly by accident? Auto-suppression only happens in Trusted mode, only for anomalies whose CO2e impact is below a rule-configured threshold, and only for suggested actions on the rule's allowlist. Every auto-resolution is a typed audit event with a reverse instruction attached, so any accidental suppression can be walked back. Most organisations run in Copilot mode for the first two reporting periods before graduating rule by rule.
How does anomaly detection interact with the review queue? The document review queue catches issues at extraction time. Anomaly detection catches issues at the ledger level after the record is written. They are complementary. A number can pass extraction confidence checks and still be a statistical outlier against the site's history; that is what the anomaly layer is for.
Can we configure the sensitivity per material or per facility? Yes. The statistical thresholds (z-score bounds, rate-of-change multipliers) and the trust orchestrator's auto-resolve thresholds are configurable per rule. A refrigerated logistics operator running R-404A with a 15 percent expected annual leak rate needs different thresholds than a data centre running a sealed chiller loop.
What is the priority score based on? Four weighted components: severity at 35 percent, CO2e impact scaled against organisational totals at 30 percent, detection confidence at 20 percent, and recency at 15 percent. The single 0 to 100 score orders the queue for the reviewer and gates auto-resolution for the Trusted-mode agent.
Carbonly pricing is per project with a $100 per month workspace minimum. If you want to see the anomaly module running against a representative sample of your own ledger, email hello@carbonly.ai.