The setup

Same file, same prompt, no follow-up coaching

The input was leverantorsfakturor-lidingo-stad-2025.csv, the openly published supplier invoice ledger for Lidingö stad. Every model received the file once, with a single instruction, and was left to decide what a procurement professional would actually need.

The interesting difference is not who was fastest. It is who understood that spend data becomes a business case only after addressability, data quality and assumptions are dealt with honestly.

The prompt, verbatim

I’m attaching the file containing spend data. As a procurement professional, I’m interested in creating a data-driven business case. Create necessary tools to analyze, interrogate, and prepare this data for my work processes. The output should be an excel file.

Speed against depth

The fast answers were fast because they skipped the hard part

Claude47 minutes
ChatGPT34m 26s
Gemini1 minute
CopilotImmediate reply
FasterRuntime (compressed scale)Slower

Time alone is not a fair comparison. Copilot and Gemini were faster because they did less. ChatGPT and Claude took longer because they added QA, reconciliation, modelling depth and — in Claude's case — reusability.

Scorecard

Nine criteria a procurement lead would actually apply

Each criterion is scored from 1 to 5. Select a row to read what was being judged.

Copilot

Procurement depth
2 out of 5
Business-case usefulness
2 out of 5
Data-quality handling
2 out of 5
Reconciliation / QA
1 out of 5
Full data preserved
2 out of 5
Excel polish
3 out of 5
Reusability
1 out of 5
Transparency of assumptions
2 out of 5
Fit for executive use
2 out of 5

ChatGPT

Procurement depth
4 out of 5
Business-case usefulness
4 out of 5
Data-quality handling
5 out of 5
Reconciliation / QA
5 out of 5
Full data preserved
5 out of 5
Excel polish
5 out of 5
Reusability
3 out of 5
Transparency of assumptions
5 out of 5
Fit for executive use
4 out of 5

Gemini

Procurement depth
3 out of 5
Business-case usefulness
3 out of 5
Data-quality handling
2 out of 5
Reconciliation / QA
2 out of 5
Full data preserved
1 out of 5
Excel polish
3 out of 5
Reusability
2 out of 5
Transparency of assumptions
3 out of 5
Fit for executive use
3 out of 5

Claude

Procurement depth
5 out of 5
Business-case usefulness
5 out of 5
Data-quality handling
5 out of 5
Reconciliation / QA
5 out of 5
Full data preserved
5 out of 5
Excel polish
4 out of 5
Reusability
5 out of 5
Transparency of assumptions
5 out of 5
Fit for executive use
5 out of 5

The results in detail

Pick a model, or compare the numbers side by side

Every output file is downloadable, and each run log is reproduced in full for anyone who wants to check the reasoning rather than take the summary on trust.

Claude

Best overall output47 minutes

Most procurement-relevant, most defensible, and closest to a reusable business-case tool.

Source rows recognised
120 508
2025 spend
SEK 1.838bn
Addressable spend
SEK 683m (37%)
Opportunity range
SEK 30.6m–82.5m

What it built

  • Config-driven cleaning pipeline
  • Procurement taxonomy
  • Analysis engine
  • Multi-tab Excel workbook
  • Interrogation CLI (query.py)
  • README and packaged toolkit

Strengths

  • Strongest analytical distinction between gross, addressable, partly addressable and non-addressable spend.
  • 100 suppliers account for 80% of spend; 1 642 tail suppliers carry 13% and roughly SEK 6.9m of invoice-processing cost.
  • 402 suppliers serve two or more förvaltningar independently, covering SEK 797m.
  • Explicit about limits: invoice data is not contract data, contract coverage is inferred, no unit price means no price benchmarking.
  • Reusable for other municipalities — configuration and scripts included.

Weaknesses

  • Slowest run of the four.
  • Excel presentation slightly less polished than ChatGPT’s workbook.
  • Deliverable is a toolkit, so it assumes some willingness to run scripts.

Best use case: A defensible procurement analysis tool, not just a one-off workbook.

What this means

The model is not the deliverable

All four models can produce a spreadsheet. Only two produced something a CFO could be shown without rework, and only one produced a tool that could be re-run next quarter on a different dataset.

The separation came from procurement judgement: knowing that invoice data is not contract data, that duplicates must be flagged rather than deleted, and that a savings range is only credible when the assumption behind it is visible.

What a genuinely good output contains

  • Addressable spend segmentation
  • Supplier, category and förvaltning analysis
  • Tail-spend and low-value invoice logic
  • Opportunity model with editable assumptions
  • Full cleaned source data
  • Data-quality and reconciliation tab
  • A clear statement of what the data can and cannot prove
  • Optional reusable pipeline for repeat analysis
Artiom Kravchenko

Shall we stay in contact?

I'm Artiom Kravchenko, and I ran this test myself. If you found it useful, join the contact network and I will let you know when the next experiment, tool or event is ready.

Join the contact network

Back to all insights