Design an AI-powered workflow tool that removes the two biggest taxes on the packaging team's time — bad data at intake and manual assembly at the end — without forcing anyone to change how they already work. The team keeps their process. The tools take the toil.
Canadian Tire is one of Canada's largest retailers, with thousands of products moving through packaging development every year. Each product's journey runs through a Line Review, a cycle where vendors, merchants, brand, and product developers all contribute data into a shared Excel template before packaging artwork can begin.
On paper, the template is the single source of truth. In practice, it's the single point of failure. There's no check on what goes into the cells, no visibility into what's missing, and no way to know who's holding things up. Every error gets caught downstream, where it's most expensive to fix.
Before the first interview, we set up a structured project vault for Claude with numbered folders that follow the life of the engagement. Raw evidence (meeting notes, transcripts, workshop photos) lived separately from what we concluded from it, so every finding could be traced back to who said it and when.
Two files did the heavy lifting: a status doc updated after every session, and an audit log that recorded each design decision with its reason. On a tight schedule this paid for itself daily. New findings slotted into place instead of getting lost in a notes. And when it was time to write the readout, the synthesis was already done.
We didn't want to design from assumptions. Over four weeks we ran a structured discovery: contextual inquiries, job shadowing, structured interviews, and synthesis sessions. All aimed at surfacing the high-friction, time-consuming parts of the process.
Week 1 was discovery and requirements gathering: deep-dive workshops with business, technical, and product stakeholders.
Weeks 2 & 3 were synthesis and validation: turning findings into a concrete plan and pressure-testing it with the team.
Week 4 was the final readout. In total: 28 sessions, 30+ interview hours, 8 divisions engaged from Packaging and Merchandising to external vendors.
The tight deadline required us to sync quite often, dedicating time each day to recap what we were hearing, what we wanted to dig into more and how our next set of interview questions might need to be changed or tweaked.
Our sessions consisted of recorded white boarding exercises with half the team that were in the Canadian Tire HQ and the other half remote. Drawing up process maps and clarifying where the source of downstream impacts seemed to be located and updating our project vault with each interview transcript to be analysed with Claude.
We created journey maps of the packagin teams process, updating as we interviewd more departments whose actions have an effect on the packaging teams deadlines. This helped us see early warning signs of the kinds of problems we may want to try to solve now vs what needed more invesment of more research to truly solve.
The Canadian Tire team already had proces maps of how the work gets completed so this also helped us validate that information.
Every department had a list of painpoints brugh about by the curretn process, some over lapping but we focused on narrowing down the pain points that affected the packaging team.
A "What-If" workshop was initially planned just to validate our learnings from the first week's interviews. Based on the team's energy, it transformed into something better — a co-design session where the team themselves generated and voted on solutions.
This mattered for two reasons. First, the people closest to the problem know where the value is: the session produced 69 potential solutions across 5 idea areas.
Second, it was a chance to teach the team how AI and agentic tools could actually help in their complex day-to-day so the final recommendation wouldn't feel like something done to them.
Two findings stood out from everything we heard. Communication loops are disproportionately expensive. 59% of Line Reviews are disrupted upstream: 44% hit missing data at the gate, 15% get late data. Each disrupted Line Review adds about 10 working days. And a single missing-data question can cost anywhere from half a day to four days when the answer needs a trip to Asia.
Manual data entry multiplies errors and delays. The same data gets retyped up to 3 times per SKU as it moves vendor → brand → developer → packaging developer. File-hunting recurs 3 times per Line Review, at about 1.5 days each.
Automating intake checks, document extraction, and content generation could shave up to 13 days off every Line Review.
The opportunity was clear: AI has matured enough to tackle poor data quality at intake, and manual assembly at the end.
The solution is a web app where coordinators, merchants, brand, PROs, and vendors each do their part supplying the information that the packaging team needs, instead of email and spreadsheets. Nine features would define it.
We wanted a visual exploration of what the solution could look like given all the pain points we had outlined.
We tried to use Claude as I had seen success in seeing how it intepreted pain points and suggested solutions, and more often than not a few features were worth integrating. So we tried instead of having it suggest improvements, using it to suggest an approach with a visual representation.
Lesson - It takes significantly more time, tokens and setting up to get Claude to get you a decent result(as of Opus 4.8). Not quite yet a helpful assistant for now especially with a tight deadline.
Taking inspiration from what we had seen while researching other helped us get an idea for how this app might come together.
We translated the workflow into concept screens for the vendor portal: a SKU list with per-SKU progress, drag-and-drop document upload where the agent sorts information into the right fields, and a persistent AI chat panel that can answer questions, give examples of good responses, and track progress to completion.
A deliberate design decision: the agent doesn't coach while a vendor is typing. It reviews the entry when the vendor signals they're ready at the "Next" action. It lets you finish your thought before responding.
We mapped every role in the Line Review: Packaging, Merchandising, External Partners, Product Development and marked exactly which steps move into the app and which stay offline. The app doesn't take over the process; it wraps the messy middle where data collection and handoffs live, and fires notifications at the moments where things stall today.
You give the agent product info and assets it gives the team the outputs they need, faster. Tech packs, SKU lists, vendor inputs, and historical Line Reviews go in.
An agent harness picks the right model and tools for each task, reads the SharePoint file, extracts template fields, matches SKU records, and sends notifications.
Out come an enriched template, flagged field gaps, proactive alerts, and a generated copy deck. And because the work happens inside a tool, measurement becomes possible for the first time.
A Simple, Maintainable Architecture.The architecture was designed to pass an enterprise IT review: a web app on Azure (Canadian Tire's existing Microsoft environment), four narrow integration points (SharePoint, Entra ID, Teams, email), and an LLM gateway with model abstraction so the client can choose which models to use for which tasks. Minimal read/write changes to existing systems. Excel stays a first-class artifact.
We modeled the value of the Packaging Assistant across adoption scenarios:
The bottom line: a build delivering $1.2M to $4.0M in net benefit over five years, with ROI between 4.1x and 7.7x and payback in as little as 10 months.
Without setting out to, we followed the BMAD Method. A framework for AI-assisted product development where each phase's documents become the context for the next. Our discovery and co-design workshop were its Analysis phase; the PRD, feature list, and wireframes were Planning; the architecture, roadmap, and risk register were Solutioning. Everything lived in a structured vault with a decision log recording every call and its reason, so every claim in the final readout traces back to a named person in a dated interview. That discipline is why four weeks of discovery ended in a costed, buildable plan rather than a deck of opinions.
The readout was delivered to Canadian Tire leadership in June 2026. And was very positively received by the client.The readout landed with a concrete plan, not just a recommendation: a 16-week path from discovery to a pilot the packaging team can run with.
We also named the risks honestly such as IT alignment, vendor access permissions, seasonal timing, each with mitigations and next steps.
What would I change? Nothing about the approach. But with more time I would have pushed the vendor portal concepts to higher fidelity and put them in front of a real vendor, and run a small prototype of the extraction step against live templates
The biggest lesson was that designing an agentic tool for people who have never worked with AI is less about the model and more about restraint. The team didn't want to be coached through every keystroke, they wanted the tool to stay out of the way and speak up at the right moment, which is why the agent reviews at "Next" rather than while you type.
Turning the validation workshop into a co-design session changed the direction of the whole engagement. The strongest ideas in the final recommendation came from the people who live in the process, and that made the plan theirs instead of ours.
What would I change? Nothing about the approach. But with more time I would have pushed the vendor portal concepts to higher fidelity and put them in front of a real vendor, and run a small prototype of the extraction step against live templates, proving the accuracy rate rather than projecting it.