Our Technology

The infrastructure behind
financial institutions’ ESG data sourcing.

Our data pipeline sources, extracts and structures ESG data from public company information. Our AI does the repetitive work, and human experts review every datapoint before delivery.

Annual reports
Sustainability reports
Public filings
Policies & codes
Company websites

01 · Sources

Everything a company officially publishes.

AI extracts and structures the data
Any company, any country, any language
Extracting from text, tables and graphs.
Unified units and currencies
Confidence score
Anomaly detection
Every figure traced to its page and line
Human in the loop ✓ Verified
Our ESG data experts review and approve every datapoint before it is delivered.
Continuous monitoring
Our platform tracks new or updated disclosures, keeping the data constantly current.

02 · Extraction & Quality

AI structures the data, analysts sign off, monitoring keeps it current.

‹ › ⚲  app.cleartraced.com
Cleartraced’s ESG Data Platform
Platform
Search and compare
API
Integrate into your systems
File export
CSV, XLSX, JSON

03 · Platform

Data is available in our platform, as file exports, or via API.

FINDING COMPANY INFORMATION

Everything a company officially publishes.

EXTRACTING, STRUCTURING & VERIFYING DATA

AI structures the data. ESG data experts sign it off. Agent monitoring keeps data current.

PLATFORM DELIVERY

Delivered via Cleartraced’s platform, offering file export or API integration, whatever fits your infrastructure.

THE CLEARTRACED DATA PIPELINE

From scattered company information to structured, auditable ESG data.

Adding a company

We only need three inputs to run our process: the company’s country, the reporting year, and its name or identifier. New companies can be requested in whichever way is easiest for you: via email, a form within our data platform, an API request, or a file over SFTP, always handled with total privacy and confidentiality.

Request · 3 inputs

Name or identifier

Volvo Cars ABor LEI · ISIN · Ticker

Country

SwedenAny jurisdiction globally

Reporting year

2025Historical data available

Finding public information

Our system scans each company’s official webpages, downloads its annual, sustainability and regulatory reports, and combs the full site for further ESG disclosures such as human-rights policies and codes of conduct, tying every document to the right entity.

SOURCING

Found: Volvo_Annual_Report_2023.pdf
Found: Volvo_Sustainability_Report_2023.pdf
Found: Volvo_Code_of_Conduct.pdf
Found: Human_Rights_Policy.pdf
Found: Supplier_Code_of_Conduct.pdf
Found: Whistleblowing_Policy.pdf
Found: Official webpage (English)
Found: Official webpage (native language)

Process reports

Our system reads PDFs, scans complex tables and charts and converts them into structured, machine-readable text, handling complex layouts and any language.

TEXT
CHART
TABLE
IMAGE

Structured JSON

{
  "text":  "…",
  "table": [],
  "chart": {},
  "image": "…"
}

Extract & structure

Structured AI prompts extract each datapoint in context, with high precision, and harmonise it to standard units.

As reported (PDF)

Scope 1312,000,000 kg
Scope 286,000,000 kg
Scope 34,120,000,000 kg
RevenueSEK 247.9B

Normalised

Scope 1312,000 tCO₂e
Scope 286,000 tCO₂e
Scope 34,120,000 tCO₂e
RevenueEUR 21.5B

Units converted (kg → tCO₂e) and currency harmonised (SEK → EUR)

Normalise & standardise

Every value is mapped into common taxonomies and regulatory frameworks before storage, enabling consistent cross-company comparison.

GHG_scope1 → ESRS E1-6 · EBA P1
energy_total → ESRS E1-5 · GRI 302
board_gender → ESRS G1 · GRI 405

Verification

A rule-based system scores every datapoint across five layers. Anything below the confidence threshold is routed to an ESG analyst; nothing is guessed or dropped. 

Automated checksOngoing

Quality score

Statistical range check
YoY consistency
Unit verification
Source-match confidence
LLM cross-validation
Flagged for expert review resolved
Expert human verificationOngoing
Cleartraced Citations review view on a laptop: extracted ESG datapoints with YoY comparison and confidence score linked to the source disclosure PDF

Store & deliver

Validated data is stored securely and delivered in real time via platform, API or file, each datapoint linked to its source and re-verified against the original PDF.

PlatformFileAPI
Cleartraced platform login screen shown on a laptop

THE TECHNOLOGY ADVANTAGE

What our technology enables that humans and legacy providers simply cannot.

ANY
company · any size
a few hundred
capped by headcount

Human teams can only cover so many names. Our pipeline has no ceiling: any company that discloses, whatever its size or market.

ANY
datapoint you need
fixed set
vendor’s menu only

No fixed list of metrics. If a company discloses it, we extract and structure it, down to the datapoints your analysis needs.

ANY
reporting language
English-first
translation gaps

AI reads disclosures natively in any language. No translation queue, no blind spots for non-English markets.

Always
current, automatically
12–24 mo
legacy vendor lag

We track each company’s release dates and monitor for new disclosures. The moment data is published, the pipeline picks it up.

FAQ

Common questions about our technology

How does Cleartraced find the public information and PDFs? +
Our pipeline automatically discovers and retrieves annual reports, sustainability disclosures, and regulatory filings from company IR pages, stock exchange databases, and regulatory repositories, across any language and jurisdiction. No manual sourcing is required.
What AI models power the extraction? +
We use several large language models, including OpenAI, Claude and Gemini, among others, choosing whichever delivers the best accuracy for each task, combined with custom-trained classification and normalisation layers. The models are continuously evaluated against a ground-truth dataset curated by our in-house ESG analysts.
How is data accuracy verified before delivery? +
Every extracted datapoint passes through a six-stage validation pipeline: document parsing, AI extraction, cross-field consistency checks, unit and boundary validation, statistical outlier detection, and a final human sign-off for low-confidence values. Each datapoint ships with its source citation and confidence score.
Can the pipeline handle non-English reports? +
Yes. Our models are trained on reports across 30+ languages including Mandarin, Japanese, German, French, Spanish, and Nordic languages. All output is standardised to English with consistent units regardless of source language.
How quickly is data updated after a company publishes a new report? +
Our monitoring system detects new disclosures within hours of publication. Extraction and validation typically complete within 24–48 hours, compared to weeks or months with manual workflows.

Built with cutting-edge technology

Anthropic OpenAI AWS Supabase Pinecone Vercel

Want to see it in action?

Book a demo and we’ll run a live extraction on any company worldwide.