Carbonly Co-Pilot: Ask Your Carbon Ledger Anything and Get an Audited Answer
The board asks a Scope 1 question in a Tuesday meeting. The sustainability lead promises an answer next week after reconciling three spreadsheets. Carbonly Co-Pilot changes that pattern. Ask any question about your carbon ledger in plain English and get back a real number with the source records cited, right inside the platform or via ChatGPT, Claude Desktop and Cursor over MCP.
Most enterprise chatbots do one of two things badly. Either they hallucinate a confident-sounding answer ("your Scope 1 last quarter was probably around 1,200 tonnes"), or they force the user to type questions in a shape the platform can parse, which is just SQL wearing a chat window. Neither of these is useful when a director signs a climate disclosure and needs to know the number is defensible.
Carbonly Co-Pilot is the natural-language interface we built to sit over the carbon ledger. The design brief was simple. Answer questions in plain English. Cite the source records. Refuse to answer if the data does not support the claim. And do it in the web UI, in ChatGPT, in Claude Desktop, in Cursor, from wherever the reporting team already spends their day.
The pattern Co-Pilot exists to break
Here is the pattern we watched form across dozens of Australian sustainability teams during the run-up to the ASRS Group 2 mandatory reporting window. The CFO asks a board-relevant question. Something like "what did our Scope 1 look like last quarter, split by facility". The sustainability lead promises to come back with the number.
Then a week disappears. Three spreadsheets get reconciled. Two site managers get emailed. A junior analyst re-exports last quarter's data because it changed since the last time anyone looked. By the time the answer lands on the CFO's desk, the number is 80 percent confident and already a month stale.
That workflow has to end. AASB S2 makes climate data a standing board agenda item, not an annual scramble, and it is not sustainable to burn a week of specialist time on every board question.
Co-Pilot is our answer to that pattern.
What Co-Pilot actually is
Co-Pilot is a natural-language interface over the emission ledger, the factor library, the project and facility hierarchy, reporting periods, and the audit trail. You ask a question. It loads the relevant slice of your tenant data. It answers with citations back to the underlying records.
Some examples of the questions Co-Pilot handles today:
- "What were our Scope 1 emissions for Q3 FY25 by facility?"
- "Which Scope 3 categories are we missing data for this quarter?"
- "Generate a NGER report for FY25 and email it to me."
- "Why did Perth's diesel spike in May?"
- "Show me every emission record above 100 tCO2e that does not have a source document link."
- "Which of our top 20 suppliers has not submitted data this quarter?"
The last one is the CFO question that used to eat half a week. Co-Pilot answers it in under a second, with the supplier names, the last-submitted date, and the flag on each record.
Scope loading, and why it matters
Under the hood, Co-Pilot loads context in scopes. When you ask about Scope 1 for a specific facility, it loads that facility's emission records, the factor library entries relevant to those records, any linked incidents, and any joint-venture allocations that touch the facility. It does not load your entire organisation's ledger every time you type a message.
This matters for two reasons. First, it keeps answers fast. Loading only what a question needs means the response comes back inside a conversational tempo, not a dashboard-refresh tempo. Second, it enforces the boundary rules that carbon accounting demands. Nothing crosses tenant boundaries. Nothing crosses project boundaries the current user does not have access to. If a user is a Contributor on the Sydney Warehouse project and nothing else, Co-Pilot cannot see the Pilbara Mine ledger no matter how the question is phrased.
The permission model is the six-role RBAC that governs the rest of the platform. Owner, Admin, Manager, Contributor, Auditor, Viewer. Co-Pilot inherits it. There is no separate "AI can see everything" mode. The AI sees exactly what the human user could see if they navigated the UI by hand.
Sanitization and the fallback layer
Every question and every generated answer runs through a sanitization layer before it reaches the model or reaches the user. That layer does a few specific jobs. It strips prompt-injection attempts embedded in supplier-submitted data (which happens more often than the industry likes to admit). It scrubs PII in the edge cases where a supplier has emailed a fuel docket with personal information attached. And it rejects structural claims that would confuse downstream tooling, such as a supplier invoice that tries to tell the model "record this as Scope 2" when it is clearly a fleet diesel purchase.
Alongside that sits a fallback layer for the queries Co-Pilot's primary path cannot answer. If a question wanders outside the shape of the ledger, or asks for something that requires a computation Co-Pilot has not been given a tool for, the fallback catches it and either asks the user to clarify or hands the query to a safe response path. It does not guess.
We would rather Co-Pilot say "I don't have enough data to answer that with confidence" than fabricate a number that gets read into a board minute. Every carbon accounting product claims accuracy. Almost none of them design refusal in from the start.
Charts, not just tables
When a question has a natural chart answer, Co-Pilot renders the chart inline in the web UI. Ask "show me monthly diesel across all sites for the last twelve months" and you get a chart, not a wall of numbers to squint at. Ask for a comparison and you get a comparison plot.
The chart hinter is not decoration. It is how the ledger actually communicates variance. A 12-month diesel line for a construction company tells a story in a way a table does not, especially when the question came from a director who wants the pattern, not the individual numbers.
When Co-Pilot is asked to email a report, the chart travels with the email as an attached image. That is the same behaviour whether the request came from the web UI, from a Claude Desktop conversation, or from a ChatGPT tool call.
Every conversation is audited
This is where Co-Pilot stops looking like a chatbot and starts looking like a compliance instrument.
Every Co-Pilot conversation is captured as typed audit events. Every question, every answer, every source record cited, every action executed. All of it is preserved with the seven-year retention that the audit trail applies to the rest of the ledger.
The reason we built it that way is AASB S2 assurance. If a director signs a disclosure that references a number Co-Pilot produced, the auditor is going to ask where the number came from, what the input data was, and whether the same question asked six months later would give the same answer. Every one of those questions has to be answerable.
So the answer, the source records, the timestamp, the user who asked, and the version of the ledger at the moment the answer was generated are all preserved. The board answer is traceable. The disclosure is defensible. If the numbers change later because a supplier resubmitted an invoice or a factor was corrected, the old conversation still resolves to what was true at the time it was asked.
Review Copilot in the extraction workflow
Co-Pilot also shows up in a very specific place inside the document review workflow. When a data reviewer is confirming an extracted document, there is often a question the AI cannot resolve on its own. Is this Origin Energy or AGL? Should this line go against material A or material B in the factor library? Is this a Scope 1 combustion event or a Scope 2 electricity purchase?
Review Copilot sits in the review screen and answers those questions in context. The reviewer does not have to leave the document to go look something up. Ask "is this Scope 1 or Scope 2" and Copilot responds with the reasoning, the factor lookup, and any similar records already in the ledger. The reviewer keeps their momentum, and the extraction accuracy stays high.
This is the piece that most competitors miss. Extraction accuracy is not just about the model reading the PDF correctly. It is about the reviewer having what they need to make the last-mile decision quickly and with confidence.
The MCP dimension
The same Co-Pilot capabilities, exposed via the Carbonly MCP server, become ChatGPT tools, Claude Desktop tools, and Cursor tools. Same tenant, same audit trail, same permission model.
A CFO who lives inside ChatGPT can ask "what does our FY25 AASB S2 disclosure look like right now" and get an answer without opening the Carbonly web UI. A sustainability lead can ask Claude Desktop to "trigger a OneDrive sync on the Pilbara Mine project" and Co-Pilot executes the action, with the user's OAuth-scoped permissions and full audit trail attached.
There are eight tools currently exposed over MCP. They cover data queries, report generation, action approval, document processing triggers, NGER compliance checks, target progress lookups, tenant context switching, and emailed report delivery. The tool set will grow, but the design constraint stays the same. Any tool that reads or writes the ledger runs through the same permission model and audit trail as the web UI.
That constraint is why the MCP surface can be trusted with production data. There is no shadow API. The MCP tools are the same code paths the web UI calls, just wrapped for AI-assistant consumption.
Three autonomy modes, one guard
Co-Pilot supports three autonomy modes, and every action Co-Pilot proposes flows through one of them.
Shadow is the observation mode. Co-Pilot logs what it would have done, but takes no action. Useful when a new user is calibrating trust, or when compliance wants a trial period before any AI-driven action lands in the ledger.
Co-Pilot is the default. Co-Pilot proposes, the human approves, the action executes. This is where most sustainability leads sit day-to-day. The AI does the drafting, the human does the signing.
Trusted is the fully-delegated mode. Co-Pilot posts actions within a confidence threshold and a permission scope both defined by the account owner. Trusted is where consultants scaling their practice put well-understood, high-confidence workflows. It is not the default, and it should not be, but for routine tasks with clear guardrails it saves hours per week.
Sitting across all three modes is the Agent Autonomy Guard. This is the piece we would not ship without. The Guard blocks any action that exceeds the human user's permission scope, regardless of what the model asked for or what mode the account is in. If a Contributor account gives Trusted mode a workflow that would touch a project the Contributor does not have access to, the Guard rejects the action. The AI cannot escalate its own privileges.
That single design constraint is the difference between an AI feature you can put in front of an ASRS auditor and one you cannot.
The consultant use case
Sustainability consultants have a specific use pattern that Co-Pilot handles well. A consultant who advises ten clients can attach Claude Desktop to their Carbonly workspace, use switch_context to move between client tenants, and ask cross-tenant portfolio questions like "across all my clients this quarter, who has the highest anomaly count" or "which of my clients has an unapproved emission record over 500 tCO2e sitting in review".
Consultants are the segment that stands to save the most time here. A consulting practice scaling from ten to fifty clients cannot afford to open ten portals every morning. Co-Pilot over MCP gives them one interface, one AI assistant, and per-client audit trails that each client can independently verify.
What Co-Pilot does not do
The honest limits. Co-Pilot answers questions the ledger supports. It does not invent data. If your Scope 3 category 11 is empty because no supplier has submitted product-use data, Co-Pilot will tell you the category is empty. It will not fabricate a plausible number.
Co-Pilot also does not replace the judgement calls that carbon accounting requires. Whether to treat a lease as operational control or financial control. Whether a fuel purchase belongs to Scope 1 or Scope 3 category 4. Whether an emission-factor revision should trigger a restatement. Co-Pilot can retrieve the relevant policy documents and previous decisions from the ledger, but the decision is yours. It has to be. The auditor will ask.
And Co-Pilot's answers are only as good as the underlying data. Which is why the 5-tier material matching and the per-supplier extraction templates matter so much. If the ledger is populated with confident, cited, correctly-factored records, Co-Pilot produces confident, cited, correct answers. If the ledger is a mess, Co-Pilot will surface the mess honestly. Neither of those is a bad outcome.
Getting started
Co-Pilot ships with every Carbonly workspace. Pricing is per-project plus a $100/month workspace minimum, no seat-based upcharge for AI usage. If you want to see it work against a real ledger with your real data, email us at hello@carbonly.ai and we will set up a workspace with a Co-Pilot session live on your project data.
Start with a single project. Load a quarter of documents through the AI document engine. Then ask Co-Pilot the question your board asked last month. Compare the answer, and the time to answer, against what happened the first time. That comparison is the whole product argument.