Arc Collective
← All case studies

Case study 02 · Custom apparel supply-chain platform

A good business, an untouchable system.

Nothing wrong with the business itself. What capped it was a system nobody could safely touch, and the thirty-two people working around that system. They came to us asking whether to rebuild it. We told them not to. Six months later that team was twenty people, not one line of the old system had been changed, and the business could take on more work than before, not less.

500
People company-wide
$28M
Annual revenue
32 → 20
Operations headcount · six months
$840K
Annual cost of that team

The company

Capped by hands, not by demand.

Made-to-order apparel, sold and manufactured through a single platform: dealers, manufacturing partners and suppliers all transacting inside the same system. Repeat accounts, real margin.

The first conversation was not about AI. It was about technical debt — eight code modules, three back offices running in parallel, new hires taking too long to become useful. They wanted an opinion on whether a rebuild was worth doing.

We took an inventory and came back with a different answer. Two generations of technology stacked on top of each other, with transaction behaviour depending on how a method had been named — the risk of touching that far outweighed the return. Meanwhile that thirty-two person operations team was costing $840K a year, fully loaded.

The old system was not the expensive part. The people working around it were.

So the engagement was redefined: leave the old system alone, build AI capability alongside it, and take the operations team out of the repetitive work. That turn got the project approved inside a week — because it removed downtime risk, which was the single biggest obstacle to a decision.

Discovery

Measure the hours before deciding what to change.

This is where most AI transformation projects die — interviews without data validation, and the whole plan ends up resting on the impression that everyone is busy. Three weeks, three lines of enquiry running at once.

Line A

Non-invasive code inventory

The true size of the estate, established without disturbing anyone: eight modules, 158 controllers, 348 server pages, 766 front-end pages, 176 distinct spreadsheet import and export methods, 153 reporting routes — and exactly two scheduled jobs in the whole system.

Line B

Shadowing and time sampling

Twelve roles observed across five consecutive working days, logged at fifteen-minute granularity, separating productive work from waiting, from context switching, from rework.

Line C

Reverse validation against data

Every interview conclusion checked against real production records. This is the difference between what we do and pure consulting. Impressions mislead. Data does not.

The finding that mattered: effective utilisation was only 68%.

Of the 6,144 nominal hours a month — 32 people × 24 days × 8 hours — only about 4,178 were landing on work that produced a business result. The missing 32% was not idleness. It was structural: 17% waiting, 9% switching between systems, 6% rework.

That changed the shape of the solution. At 95% utilisation, AI's only route is to replace the work itself, and the ceiling is low. At 68%, there are two sources of gain — AI replacing work, and process continuity removing the waiting, the switching and the rework.

Findings

Every one confirmed in code or in data.

Not interview impressions. Each of these points at a specific piece of live code or a specific set of production records.

01

The queue rejected orders on people's behalf

Live code auto-rejected any order sitting unreviewed for more than fourteen days. On the numbers, 143 orders a month were being lost that way.

02

One person, five systems

A single transaction required checks across several back offices and consoles. Measured during shadowing: eleven switches per person per day.

03

Export was the workflow

Behind those 176 export methods sat an entire offline routine — open the page, set the filters, export, then rework it in a spreadsheet. And none of those endpoints had usage instrumentation. Nobody knew which reports were still being read.

04

Five parallel definitions of the same statistics

Sales, finance, payments, after-sales and search each used their own combination of order states. Cross-department reconciliation regularly failed to agree. The numbers were not wrong. The definitions were different.

Approach

Grade the work first, then decide what can be automated.

We did not start by asking what AI could do. We started by grading every operational task by how automatable it actually was. That framework set the ceiling on every reduction figure that follows.

Grade A · 61% of the work

Fully automatable

Deterministic rules and finite enumerations — credit and status checks inside order review, SKU matrix generation, column mapping and unit normalisation. Criteria are enumerable and results recomputable. AI takes these over; people sample-check.

Grade B · 24%

Human in the loop

AI locates the problem and presents the evidence and a recommendation; a person makes the call. Compressible by 30–60%, but never to zero, because accountability has to sit with a person.

Grade C · 15%

AI cannot remove it

Phoning a customer about a material shortage, agreeing price and process across departments, photographing physical product. The direct reason the reduction stops at 69% rather than going further.

The engineering red line: not one line of legacy code.

Every AI capability runs as an independent side-car service, reads from read-only data sources, and performs every write through existing interfaces. It does not bypass existing authorisation and it does not write to state tables directly. That red line got the design through the client's technical committee on the first pass.

One job, start to finish

A customer's spreadsheet.

Taken from the system running today. The example happens to be an apparel order sheet. Substitute an insurance claim schedule or a bill of materials and the shape does not change.

inbound · customer email attachment
“Colour / Size / Qty / Fabric spec” — four column headers, written in this customer's own vocabulary.
What the system needs is a colour code, a specification code, a quantity and a process attribute. Four fields, and not one of them matches.
01

Reads the headers

Works out which system field each customer column corresponds to — not against a fixed mapping table, but by meaning, so that “fabric spec”, “material” and “cloth” all resolve to the same field.

02

Normalises the values

Colours, sizes and units aligned against finite enumerations. “Navy”, “dark blue” and the customer's own shade name converge on one colour code; “dozen” and “12 pcs” on one unit.

03

Flags the problem rows

Rows that cannot be resolved are pulled out with the reason stated. Nobody reads the whole sheet. They handle the rows that were flagged.

04

Validates before writing

Target data structure validated before any write. Failed validation returns the row rather than writing it. The worst case is an import that did not happen, not bad data that did.

38 minutes a sheet at a 46% first-time success rate, before. Nine minutes at 94%, after.

Handover

Trust is bought with data, not promised.

One principle we held to: nothing AI decides goes into the database on day one. Order review directly affects orders and revenue. If the client is going to trust it, the trust has to come from somewhere.

  1. Shadow modeAI decided every order but wrote nothing, with each decision compared against the human outcome. Over two weeks, more than 3,100 orders compared. Only once agreement held steady at 98.6% did we move on. Operations workload during those two weeks: unchanged.
  2. Controlled releaseAutomatic clearance opened first to existing customers, existing styles and order values inside the historical range. Starting clearance rate 71%.
  3. Full runningClearance boundaries widened week by week. By the second month, 88% clearing automatically at 99.3% sample-check accuracy, with the sample rate reduced from 10% to 5%.

Results

Two months of comparison running.

Measured item by item against the pre-change baseline. Every figure below comes out of the measurement layer built into the system. None of it was estimated afterwards.

MeasureBaselineMonth 1Month 2Change
Operations headcount322620−37.5%
Effective workload (hours/month)4,1783,1802,647−36.6%
Utilisation per person68%76%82%+14pp
Average review handling time4.2 h1.1 h0.6 h−86%
Orders clearing automatically0%71%88%—
Sample-check accuracy on cleared orders—98.6%99.3%—
Orders auto-rejected on timeout143349−94%
Spreadsheet import first-time success46%88%94%+48pp
Time per customer spreadsheet38 min14 min9 min−76%
SKU time to listing2.7 days1.1 days0.6 days−78%
Data entry time per style52 min21 min14 min−73%
System switches per person per day1164−64%
The check that usually gets skipped: nobody ended up working harder.

Before the change, each person was carrying 131 hours a month. After it, across twenty people, 132. Effectively identical. The reduction did not come from pushing the people who stayed — it came from work that should never have existed being removed. That is the reason this holds as a steady state rather than a sprint.

And the business can take on more, not less. Orders auto-rejected on timeout fell from 143 a month to 9 — that is 134 orders a month kept instead of lost, straight onto the revenue line. Review handling time fell from 4.2 hours to 0.6, and in this business the supplier who answers first usually wins the order. After twenty people absorb the entire workload there is still 15.9% of capacity in hand: room held for growth, not room sitting idle.

The attribution has a control. Order volume moved within ±4% across the comparison period, and the four functions left untouched — reconciliation, after-sales, reporting and customer service — showed no material movement over the same window. So the change attaches cleanly to the process work.

Whether this transfers to you

Do not start with what AI can do. Start with whether you are this shape.

01

Is your output equal to what your team can process by hand in a day?

The important one. If that sentence is true, your ceiling is in your headcount, not in your market.

02

Is there a system you cannot touch?

Not because nobody wants to — because touching it breaks things.

03

Is a large share of the work enumerable judgement?

Criteria that can be written down and results that can be recomputed — reviewing, importing, entering.

Three yeses and it is the same shape. The first move is not installing anything — it is measuring your own baseline. Every reduction figure in this case study has as its denominator hours measured inside that company, not an industry average.

Client name available under a non-disclosure agreement.

← Back to case studies