Before we agree to "clean up the data first," I'd like to know what done looks like. Clean enough for what? We need to know the job we're asking AI to do, otherwise a small pilot can end up responsible for everything we've been meaning to fix for years. That's quite a lot to ask of a pilot.
And, I get the temptation. If you've been asking for a repair for years and there's finally some interest in paying for it, of course you want to get it done. The work may well deserve the money. I'd just want us to be clear about how much of it this pilot actually needs.
In How to Adopt AI Without Rebuilding Your Business, we gave an agent a specific job before deciding what it needed. Let's do the same with the data. We might find that we can begin with less cleanup than we expected, or that something we thought was a minor problem really does need fixing first.
What would the agent get wrong?
Let's go back to the customer order from the first post. We're using a hypothetical company that sells and ships physical products, and the customer wants everything delivered together. Some items are out of stock, so someone needs to decide whether to hold the shipment or ask the customer about sending part of it now.
Order Researcher is the agent we proposed to help gather that information. It brings together the order, stock information, and customer instructions, with links to the records. A person checks its work and decides what happens next. We also considered Order Resolver, which would get permission to change orders within agreed rules. We're preparing for Order Researcher first.
What happens if Order Researcher misses the request to deliver everything together? It could get the stock figures right and still leave the person thinking a partial shipment is fine. More accurate inventory won't help with that. We have to find the delivery request and connect it to this order.
Now we've got something specific to fix. Where is that request recorded, and how do we know which order it belongs to? If the answer depends on somebody recognizing the customer's name, we need to work out how Order Researcher will make a reliable match too.
Compare that with duplicate contact records elsewhere in the company. Do they affect this order, its instructions, or who can see the research? Check before deciding. If they don't, I'd keep them on the cleanup list without making this pilot wait for them.
For each proposed repair, we should be able to explain what Order Researcher would get wrong, miss, or be unable to show the person without it. That's a useful question to keep asking as the list grows.
Ask how the people know
Before we send the whole problem over to the data team, ask someone who handles orders to work through this one with you. How do they know the request applies to this order? What makes them check with the salesperson instead of trusting what's in the system?
There was probably a good reason to ask a colleague. Maybe the software couldn't handle the request and a conversation got the order moving. Fair enough. But, if we're handing some of that work to an agent, we need to understand what the conversation settled. The order record alone may not tell us.
We might need to start recording delivery requests against an order number, or give Order Researcher access to a source the team already uses. We might also decide that some requests still need a person to investigate. All reasonable options, but they call for different work. Calling the whole thing "data cleanup" doesn't tell the team which one we're asking for.
And, what if the email says to deliver everything together but the order record allows partial shipments? Who gets to decide which one applies? The technical team can make both records available. We still need the people responsible for the order to agree on how to resolve the disagreement.
I'd be happy for Order Researcher to show the conflict and hand it to a person, as long as it reliably catches the disagreement and brings the relevant records along. That saves someone some hunting. Choosing an instruction on its own would give Order Researcher a job we haven't agreed to hand over.
Check the permissions while we're at it. Order Researcher needs access to its sources, and the person reading its answer needs permission to see what's in them. If that doesn't work, we have an access problem to resolve or a use we need to change. Putting the information in an AI answer doesn't give everyone permission to read it.
Yesterday's answer may be the wrong answer today
How current does the stock information need to be? I'd answer that by looking at the decision we're trying to help someone make.
Suppose the warehouse sends inventory figures to the order system overnight. Order Researcher could quote those figures perfectly and show when they were updated. Useful, but can we promise the customer a shipment this afternoon based on them? Other orders may have used some of that stock since then.
We need a check that's current enough for the promise we're making. Maybe Order Researcher can look up availability in the warehouse system. Maybe a person keeps doing that check during the pilot. Either could work. Just count that person's time when we're judging the benefit; asking them to repeat most of the research doesn't save them much.
We could instead use the overnight figures to ask Order Researcher about yesterday's backlog. That might require less preparation, but we'd be choosing a different job. We'd need to agree that the historical view is useful and make sure people understand its limits.
So when we say "the data is accurate," let's finish the thought. Is it accurate enough to explain what happened yesterday, or to decide what we can promise today? The answer can change how much preparation we need.
There's a similar trap when a source is unavailable. "I couldn't check the delivery instructions" is a very different answer from "the customer didn't give any." Order Researcher needs to report the gap, and we need a way to verify which sources it actually reached. Otherwise the answer can look complete simply because the missing information never made it in.
Can we begin with a smaller repair?
Maybe we can make Order Researcher useful without resolving every uncertain customer match. If some orders have reliable order numbers connecting the records we need, we could begin there and leave the others with the team.
But, we have to be able to tell which orders belong in that group. Writing "we'll exclude those cases" in the plan doesn't do the excluding. Order Researcher needs to catch a missing or uncertain match and send it to someone who has the information and time to investigate.
Then ask whether the remaining job is worth doing. How much useful research can Order Researcher complete within that limit? How much work still comes back to a person? A narrow pilot can teach us something, but I'd want us to be honest if we've narrowed it past the point of helping anyone.
If every useful version of Order Researcher needs the same missing connection, I'd take that finding back to leadership. We can explain what fixing it would let us do and decide whether to fund it or start somewhere else. We've learned something more useful than "the company isn't ready for AI."
The broader cleanup can still go ahead. It may have a perfectly good business case of its own, and we should give it one. This pilot doesn't have to justify all of it.
Knowing when we've done enough
Before the cleanup gets underway, agree on what would make Order Researcher useful to the team. Can it find the right instructions for the order and show stock information the person knows how to use? Can they see where the answer came from and tell when something still needs investigation? After checking the research, are they getting to a sound decision with less effort?
Try that on the work the team actually handles, including difficult orders and cases that weren't used to improve the agent. Change a delivery instruction and see whether Order Researcher finds the change. Make a required source unavailable and check whether it reports the gap. Look at what the answer leaves out; it can contain nothing false and still miss the fact that matters most.
Use what we learn to decide whether to start using Order Researcher for the agreed cases, fix something else, or stop. And, decide who's looking after the connections and definitions once the pilot team moves on. They won't stay useful without attention, and the team still needs a way to do the research when a source fails.
When that job works well enough to help, we can stop adding to this pilot's cleanup list. Someone still needs to own the remaining problems and decide when to tackle them. In the meantime, people can use what we've finished.
If we later want Order Resolver to change orders, we reopen the preparation question. The person reviewing Order Researcher's output may currently resolve a disagreement or confirm stock before acting. Order Resolver needs a tested way to handle those checks, or to leave the affected orders with a person. A successful research pilot tells us something about the data, but its success includes the human work we kept in the process.
That's what I'd want to come away with. We know which job we're making possible, why each repair belongs, and what the people will still need to do. We can decide whether it's worth the work, and we have a way to recognize when we're ready to give it a try.