Four AI models, one spend file: which one produced a usable business case?
Claude, ChatGPT, Copilot and Gemini were given the same public invoice ledger from Lidingö stad — 120 508 rows, SEK 1.838bn of 2025 spend — and the same prompt. Each was asked to prepare the data for procurement work and return an Excel file. The outputs were not close.
#1
Claude
Best overall output
Most procurement-relevant, most defensible, and closest to a reusable business-case tool.
See the detail#2
ChatGPT
Best quick usable workbook
Strong professional workbook, full ledger preserved, good QA and visual review — but less reusable than Claude.
See the detail#3
Gemini
Weakest final result
Did the analysis and eventually produced an Excel file, but the workbook lacked depth, QA and procurement logic.
See the detail#4
Copilot
Fastest simple output
Created a basic workbook quickly, but the analysis was shallow and the final file is a dashboard CSV rather than a complete Excel model.
See the detailThe setup
Same file, same prompt, no follow-up coaching
The input was leverantorsfakturor-lidingo-stad-2025.csv, the openly published supplier invoice ledger for Lidingö stad. Every model received the file once, with a single instruction, and was left to decide what a procurement professional would actually need.
The interesting difference is not who was fastest. It is who understood that spend data becomes a business case only after addressability, data quality and assumptions are dealt with honestly.
The prompt, verbatim
“I’m attaching the file containing spend data. As a procurement professional, I’m interested in creating a data-driven business case. Create necessary tools to analyze, interrogate, and prepare this data for my work processes. The output should be an excel file.”
Speed against depth
The fast answers were fast because they skipped the hard part
Time alone is not a fair comparison. Copilot and Gemini were faster because they did less. ChatGPT and Claude took longer because they added QA, reconciliation, modelling depth and — in Claude's case — reusability.
Scorecard
Nine criteria a procurement lead would actually apply
Each criterion is scored from 1 to 5. Select a row to read what was being judged.
Copilot
- Procurement depth
- 2 out of 5
- Business-case usefulness
- 2 out of 5
- Data-quality handling
- 2 out of 5
- Reconciliation / QA
- 1 out of 5
- Full data preserved
- 2 out of 5
- Excel polish
- 3 out of 5
- Reusability
- 1 out of 5
- Transparency of assumptions
- 2 out of 5
- Fit for executive use
- 2 out of 5
ChatGPT
- Procurement depth
- 4 out of 5
- Business-case usefulness
- 4 out of 5
- Data-quality handling
- 5 out of 5
- Reconciliation / QA
- 5 out of 5
- Full data preserved
- 5 out of 5
- Excel polish
- 5 out of 5
- Reusability
- 3 out of 5
- Transparency of assumptions
- 5 out of 5
- Fit for executive use
- 4 out of 5
Gemini
- Procurement depth
- 3 out of 5
- Business-case usefulness
- 3 out of 5
- Data-quality handling
- 2 out of 5
- Reconciliation / QA
- 2 out of 5
- Full data preserved
- 1 out of 5
- Excel polish
- 3 out of 5
- Reusability
- 2 out of 5
- Transparency of assumptions
- 3 out of 5
- Fit for executive use
- 3 out of 5
Claude
- Procurement depth
- 5 out of 5
- Business-case usefulness
- 5 out of 5
- Data-quality handling
- 5 out of 5
- Reconciliation / QA
- 5 out of 5
- Full data preserved
- 5 out of 5
- Excel polish
- 4 out of 5
- Reusability
- 5 out of 5
- Transparency of assumptions
- 5 out of 5
- Fit for executive use
- 5 out of 5
The results in detail
Pick a model, or compare the numbers side by side
Every output file is downloadable, and each run log is reproduced in full for anyone who wants to check the reasoning rather than take the summary on trust.
Claude
Best overall output47 minutesMost procurement-relevant, most defensible, and closest to a reusable business-case tool.
- Source rows recognised
- 120 508
- 2025 spend
- SEK 1.838bn
- Addressable spend
- SEK 683m (37%)
- Opportunity range
- SEK 30.6m–82.5m
What it built
- Config-driven cleaning pipeline
- Procurement taxonomy
- Analysis engine
- Multi-tab Excel workbook
- Interrogation CLI (query.py)
- README and packaged toolkit
Strengths
- Strongest analytical distinction between gross, addressable, partly addressable and non-addressable spend.
- 100 suppliers account for 80% of spend; 1 642 tail suppliers carry 13% and roughly SEK 6.9m of invoice-processing cost.
- 402 suppliers serve two or more förvaltningar independently, covering SEK 797m.
- Explicit about limits: invoice data is not contract data, contract coverage is inferred, no unit price means no price benchmarking.
- Reusable for other municipalities — configuration and scripts included.
Weaknesses
- Slowest run of the four.
- Excel presentation slightly less polished than ChatGPT’s workbook.
- Deliverable is a toolkit, so it assumes some willingness to run scripts.
Best use case: A defensible procurement analysis tool, not just a one-off workbook.
What this means
The model is not the deliverable
All four models can produce a spreadsheet. Only two produced something a CFO could be shown without rework, and only one produced a tool that could be re-run next quarter on a different dataset.
The separation came from procurement judgement: knowing that invoice data is not contract data, that duplicates must be flagged rather than deleted, and that a savings range is only credible when the assumption behind it is visible.
What a genuinely good output contains
- Addressable spend segmentation
- Supplier, category and förvaltning analysis
- Tail-spend and low-value invoice logic
- Opportunity model with editable assumptions
- Full cleaned source data
- Data-quality and reconciliation tab
- A clear statement of what the data can and cannot prove
- Optional reusable pipeline for repeat analysis

Shall we stay in contact?
I'm Artiom Kravchenko, and I ran this test myself. If you found it useful, join the contact network and I will let you know when the next experiment, tool or event is ready.
