Skip to content
KaptionCelsa
  • Case study
  • Internal audit
  • Manufacturing

How Celsa Group piloted AI-powered control testing with Kaption

Audit functions worldwide face a widening gap between increasing control volumes and limited team capacity. The result is compromised risk coverage, heavy reliance on outsourcing, and delayed insight for improvement.

To address these challenges, Celsa Group, a leading European manufacturing group, partnered with Kaption to create MAAT (Multi-Agent Audit Technology). A five-month pilot across 60+ manual and IT general controls under International Financial Reporting Standards proved that a human-on-the-loop approach lets AI expand coverage while keeping expert auditors firmly in control.

Client
Celsa Group
Headquarters
Barcelona, Spain
Sector
Steel and manufacturing
Employees
c.5,000
Legal entities
30+
Pilot scope
60 ICFR / SCIIF controls
Galvanised steel roof trusses against a pale sky
Steel operationsCelsa Group

01 The problem

An industry-wide constraint

Many enterprises face the same assurance gap: the volume of controls their audit function is expected to test has grown well beyond what the in-house team can realistically cover.

The result is heavy reliance on outsourcing, slow and costly insight into where controls are failing, and risk coverage that falls short of what regulators and boards expect. Fewer than half of Chief Audit Executives feel confident in their team's ability to assure rapidly expanding risk and compliance demands.

For most multi-entity enterprises, full-population control testing is simply not achievable with existing resources: too many controls, too many of them informal or undocumented, too few auditors, and testing cycles that stretch over months rather than days.

Controls to test versus in-house capacity
Controls to testIn-house capacityUntested riskTodayTimeVolume
  • Controls to test
  • In-house capacity
  • Untested risk

Source: Gartner, “Internal Auditors to Focus on Cybersecurity, Data Governance and Regulatory Compliance in 2026”, press release, 13 November 2025.

02 The challenge

Celsa, a textbook case

Celsa Group is a textbook example of this industry-wide challenge: a control universe of more than 800 distinct controls, generating over 4,500 testing instances a year, managed by an internal audit team of fewer than five FTEs.

800+

distinct controls

4,500+

testing instances a year

<5

internal audit FTEs

Operating across a complex, multi-entity group, the gap between what the team could realistically cover and what the organisation required had become significant, and was widening.

As an initial response, Celsa relied on an outsourced testing provider. A single engagement covering approximately 170 controls took three months to complete. At that pace, full-population coverage was operationally and financially out of reach.

The team had already solved coordination with more than 150 control owners by adopting a GRC platform to centralise risk and control data. Testing itself remained entirely manual, which limited the value the platform could deliver and left the core capacity problem unsolved.

As the business evolved and stakeholder demands intensified, Celsa faced a compounding constraint: more coverage, more frequently, with greater rigour, and no scalable way to deliver it.

03 The solution

Introducing MAAT

Kaption worked closely with Celsa's internal audit team to develop MAAT, an AI system that automates control testing at scale and mirrors the way Celsa's auditors really work.

Rather than a generic deployment, MAAT was configured against Celsa's own control standards, audit methodology and evidence evaluation criteria, and also consolidates the expectations of external auditors. It ingests real audit evidence, analyses it against each control's requirements, and delivers a clear assessment of control effectiveness with full reasoning and source-level citations, documenting every step.

Every verdict is logged and traceable, giving audit leadership complete visibility into how each conclusion was reached. Because audit is judgement-based and relies on experience, MAAT was designed and tuned with Celsa's experienced auditors to produce justifiable, defensible and fully traceable reviews.

  1. 01 Ingest

    Diverse evidence

    Structured and unstructured files, such as PDFs, spreadsheets and screenshots, are ingested by the model.

  2. 02 Analyse

    Multi-agent reasoning

    Specialist agents evaluate evidence against Celsa's control standards and methodology.

  3. 03 Cite

    Source-level citation

    Every conclusion links back to the specific evidence and rule that produced it.

  4. 04 Verdict

    Justified outcome

    Effective, Ineffective or Needs Review. Fully logged, traceable and defensible.

04 The approach

Human on the loop

MAAT is built on a human-on-the-loop principle. Unlike human-in-the-loop systems, where the AI stops at every turn to wait for manual intervention, human on the loop shifts the auditor's role from execution to oversight.

That distinction matters. The goal is not an autonomous system that blindly signs off on controls, nor a tool that needs hand-holding at every step. The AI tests independently, at a scale previously impossible without large teams. The expert stays on the loop, supervising the entire population, reviewing the evidence the AI has compiled, and confirming final conclusions rather than performing every step by hand.

A set of safeguards strengthens the results in line with external stakeholders' expectations. MAAT is designed to know the limits of its own confidence: it gives a firm opinion when the evidence supports one, and defers to the auditor whenever it does not.

When confident

It states an opinion

Where the evidence clearly supports a conclusion, MAAT issues a definitive verdict, fully cited and traceable to source.

When uncertain

It suggests, not asserts

Where confidence is lower, MAAT offers a recommendation and flags the item for human review rather than guessing. This keeps hallucination risk low.

Safeguard 01

Evidence-bound reasoning

Every conclusion is grounded in the evidence actually reviewed and linked back to the specific source and rule that produced it.

Safeguard 02

Bias toward caution

Lower-confidence cases are routed to Needs Review rather than passed through, so the system errs on the side of caution, never leniency.

Safeguard 03

Full traceability

Every verdict is logged and auditable, giving leadership complete visibility into how each conclusion was reached, and the ability to challenge it.

Safeguard 04

Expert sign-off

The auditor remains the final authority. MAAT prepares and justifies; the human reviews, confirms and owns the conclusion.

“Far beyond anything similar I've ever seen in the market.”
Edgar Garcia, Head of Internal Audit & Risk Control, Celsa

05 The results

Pilot outcome

The results validated the approach. All control verdicts met expectations, and the system behaved exactly as a high-stakes audit environment requires.

In 85% of cases, MAAT's reasoning not only reached the right verdict, it matched the quality and structure of what an experienced auditor would produce. All outputs were grounded in real evidence, clearly reasoned and fully traceable to source.

MAAT produced fewer than 10% false negatives across the entire pilot. Where confidence was lower, the system flagged the assessment as Needs Review rather than passing it through.

On average, MAAT completed the assessment of a control in under two minutes, with a fully cited explanation ready for auditor review.

Reasoning alignment

85%

MAAT against the auditor baseline85 / 100

False negatives

<10%

  • Verdict confirmed
  • Flagged for review

Time per control, on average

<2minutes

With fully cited reasoning

Pilot coverage

50controls

  • Procure to pay
  • Order to cash
  • Payroll
  • Inventory
  • Treasury
  • Valuation
“Auditors rarely agree word for word. 85% alignment on reasoning is more consistent than most human hires.”
Marc Cano, Internal Audit & Risk Control Team Lead, Celsa

06 The impact

Benefits

When AI handles first-line control testing at speed and to standard, the coverage gap closes. A fully tested control matrix across the full population, not just a sample, becomes operationally achievable for the first time.

  1. 01

    Higher-frequency testing

    Controls can be tested more often, including those currently sampled only once a year.

  2. 02

    Broader entity coverage

    Entities previously out of scope because of capacity limits can be brought into the testing population.

  3. 03

    Auditor capacity unlocked

    Time previously spent on first-line testing is redirected to higher-value, judgement-led work.

07 Next step

What's next

The pilot results were strong. The next phase takes MAAT from pilot to group-wide rollout.

MAAT's natural evolution expands coverage in two ways: moving from testing to advising, by suggesting to control owners how their controls could improve; and extending to other domains such as tax, cybersecurity, operational and safety controls across the Group.

From c.60 controls piloted to a full-group rollout across finance, tax and operational controls.

1,000+control tests a year

Ready to see governance at the speed of growth?

Book a demo