Skip to content

Insights

Nine Agents Inside a Consulting Firm: A Build Teardown

What problem did the consulting firm need to solve?

Senior leaders at a firm this size live on signal. An executive changes roles, a competitor publishes, a market moves, and the first firm in the room wins the work. The research capability existed. The problem was how it arrived: documents on request, days late, rebuilt from scratch for every leader who asked.

Worse, every new request spawned another standalone tool with its own rules, and none of them made the next one cheaper. When I traced how intelligence moved from a leader's question to a delivered answer, the same failures kept appearing:

  • Multiple standalone tool requests overlapped in sources, audiences, and delivery cadence, and every one was being built from zero.
  • Useful logic lived inside individual builds, so an improvement for one team never reached another.
  • New requests had no shared intake, no priority order, and no operating owner.
  • One executive's preferences had hard-coded themselves into results meant to serve different sectors and roles.
  • Nobody could say what a single briefing cost to produce, so nobody could decide what deserved to scale.
  • Outputs scattered across tools and inboxes, so leadership never saw the system, only fragments of it.

The visibility gap was real enough that months in, a senior sponsor could still ask me how many executives, companies, and news sources we were tracking for him. If the person paying for the system has to ask that, the system is invisible. That question is why every agent now lives on one status page the client can open any time: what it does, its current state, and its blockers.

The firm did not need another research document. It needed its recurring leadership questions turned into agents that answer them every day.

Why was one large model with one big prompt the wrong design?

The default move right now is one big model, one enormous prompt, every instruction stuffed in. I ruled that out early, for four reasons that all showed up in the audit.

Preferences bleed. The most damaging pattern I found was one executive's preferences hard-coded into results meant for different sectors and roles. That is exactly what a single shared prompt produces at scale: every accommodation for one reader silently becomes the default for all of them. The fix has to be structural. Shared research behavior lives in one place, per-leader tailoring lives in another, and there is a hard wall between them.

You cannot price a blob. Leadership wanted to compare intelligence work the way they compare business units: what does this cost to run, and is it earning its keep? One monolithic assistant has one bill. Nine focused agents each report the cost of every run, so the firm scales the ones that earn it.

Improvements need a propagation path. With standalone builds, a fix for one team dies inside that build. With one giant model, a fix for one team changes output for every team whether they wanted it or not. Focused agents with shared research rules solve both: shared behavior improves once for everyone, and nothing else moves.

Combining findings requires discrete outputs. A senior reader wants one coherent briefing, not six tool outputs. To combine findings and remove repeats, you need agents that return separate, comparable results. One detected change can then feed every team that cares about it without being collected twice.

What do the nine AI agents actually do?

Not categories. The actual work each of Zoro's nine agents does:

  1. Executive change detection. Board appointments and executive moves show up in securities filings before they show up on LinkedIn. This agent watches the specific filing type that carries them across 76 tracked companies and alerts the day they land. The first two alerts reached the client the same day the feed went live.
  2. Alumni movement. Just over 1,000 firm alumni monitored for job changes. Each cycle lands as one summary: the move, the partner who owns the relationship, and a suggested outreach note. The firm later asked to add roughly 1,400 more names, which is what expansion looks like when a system works.
  3. Earnings release watch. Two dozen companies the team actually covers. Each release is processed the next morning with a one-line read on the stock reaction, and beats and misses formatted so a partner can scan the quarter in seconds.
  4. Financial action and major news. 44 companies across nine event categories: funding, acquisitions, restructurings, and the rest of the moves that open a sales conversation.
  5. Competitor mentions. This started as a standalone agent and stopped being one. Mentions of the firm's four closest competitors are now flagged automatically inside the major-news pipeline. An agent that deletes itself into shared infrastructure is the platform argument in one sentence.
  6. Industry news curation. A whitelist of 55 approved sources, scanned continuously, so a leader reads a briefing instead of a link list.
  7. Publication mapping. Classifies the firm's own research output so it can be matched and routed instead of sitting in a library.
  8. Client article matching. Runs downstream of the mapper: which piece of the firm's thinking should reach which client, recommended per relationship.
  9. Organization mapper. Type a company name, get the executive structure assembled from public data. This one also produced the worst failure of the build, covered below.

There is also an event planner that batch-processes conference attendee lists. For one legal-industry conference week, it took a list of 2,300 attendees across roughly 1,000 companies and returned each one enriched with revenue, organizational grouping, and a likelihood-to-buy score. A senior partner used that output for real dinner invites. One run came in as an urgent request for an event the following night and went out the same day.

What does the weekly leadership brief contain?

Zoro's agents feed one weekly brief for each senior sponsor: job and board changes, earnings, key client and industry news, the firm's own thought leadership, and competitor publications, combined and deduplicated into a single email. A typical issue carries 15 to 20 scored signals. The densest single brief carried 23. The cadence held at weekly through the engagement, and the sponsor asked me to keep it running after the engagement closed, twice.

What intake rule let nine agents ship in weeks?

An earlier tool Zoro built for this client took more than six revision cycles, almost all caused by requirements surfacing late. After that, I made a context brief mandatory before any new agent: business context, voice calibration, data requirements, guardrails, and delivery specs, written down before a line of work begins. The client's team answered with a requirements document per agent within two days.

That one process change is why nine agents shipped in weeks while the earlier single tool took six-plus rounds. If you take one operational lesson from this teardown, take that one.

What does the agent research actually cost to run?

Nobody publishes this part, so here is what Zoro's research platform costs to run.

  • Buy versus build. I priced a commercial signal platform that tracks 100+ corporate event types across 10,000+ sources. The realistic annual cost penciled into the $50K to $200K range. The specific signal that mattered most, executive appointments in securities filings, comes from a $55 per month filing API. We built around feeds like that.
  • Deep research is the expensive part. The event enrichment pipeline ran five research passes per company at roughly a dollar per company per pass. A full conference run penciled out at $4,000 to $5,000. I benchmarked six alternative approaches and cut the cost per run by 85 to 90 percent.
  • Cheap has a coverage price. The obvious cost fix was a model roughly 40x cheaper per call. On the research dimension that mattered most, its coverage dropped to about 20 percent. We kept the expensive model where depth pays and the cheap one where it does not.
  • Profile matching got 90 percent cheaper. Switching the people-enrichment step from a per-record enrichment API to a search API cut that line by 92 to 94 percent at the same quality.

This is why per-run cost reporting is a platform feature and not an accounting afterthought. Every one of these decisions came from being able to see what a run cost.

Which decisions stay with people instead of the agents?

Zoro's platform prepares the research and the view. It does not pretend to make partner-level calls. Senior leaders still own priorities, sensitive interpretation, relationships, and the decision that follows every briefing.

And the humans at this level check the machine. A senior partner recounted an attendee file line by line and challenged my company total within a day; the audited count is the one in this piece. Another pass flagged a seniority score as too generous because a manager-level title had been rated like an executive. That scrutiny is not friction, it is the quality system. The agents got better because senior people treated the outputs as claims to verify, not answers to accept.

The boundary is operational, not decorative:

  • A research operator can hold any briefing for another pass and records why a finding was removed.
  • A division manager decides whether a new request becomes a platform change or stays a one-time research task, and approves every change to shared behavior.
  • A senior leader marks findings useful or irrelevant, asks for deeper research, and submits new requests. They never operate the agents underneath.

What broke during the build?

Every honest teardown has this section. Mine has four entries.

The organization mapper shipped wrong. Early versions aggressively tagged anyone with "Director" in their title as C-suite. One company came back with 725 executives mapped and 317 of them marked C-suite, which no company on earth has. On another, a senior partner who knew the leadership team personally told me the output was completely wrong. Result counts also swung from 3 people at one company to 300 at another with no explanation a reader could see. The fixes took two rounds of structural work on title classification and source verification, and the same partner who called it wrong signed off on the rebuilt version.

Discovery was the gap, not scoring. Early briefs did not compel action, and the sponsor said so directly. The instinct is to tune the scoring. The audit said otherwise: one issue carried only 8 articles drawn from 6 of the 55 whitelisted sources. The problem was that the system was not finding enough to score in the first place. The week a story with the firm's own name in the headline went around the industry, the brief missed it. We rebuilt discovery breadth before touching scoring weights.

Data errors are the trust tax. One brief attached an executive to a company she did not work for. A signal window drifted so items outside the stated date range appeared in a dated brief. Each one of these costs more reader trust than ten good signals earn back, which is why the operator review pass exists.

The first urgent run matched 83 percent. On a time-sensitive event run, 83 percent of attendees matched to enriched company data, and the client flagged it as low. They were right.

The surprise was the expansion channel. I did not sell this platform to the second, third, or fourth team. Leaders showed it to leaders. A sponsor put the tools in front of a large internal meeting within weeks of kickoff, the platform was demoed internally more than once after that, and colleagues from other teams asked for access just to explore the systems. Internal demand is the one adoption curve you cannot fake.

What results did the nine agents produce?

  • First working system live in 6 days from kickoff at the consulting firm.
  • 1 team to 4+ teams, through internal referrals, not sales.
  • 9 production AI agents handed over with full documentation.
  • 2,400+ executives and companies monitored across the firm's approved sources, growing monthly.
  • 23 signals combined into a single brief.
  • $5M+ in pipeline influenced by business development conversations sourced from platform briefings.
  • Every run reports its own cost.

Client identity, people, source lists, outputs, and commercial terms stay withheld. The anonymized record is public: the full case study at zorocorp.com/case-studies/strategic-intelligence, with the sanitized interactive walkthrough linked from it.

What did the handover include, and what happened after?

The firm received nine documented production agents from Zoro, plus the platform they run on: source code, deployment guides, configuration templates, and data model documentation. One intake for new intelligence needs. Shared behavior kept separate from personal preferences. Per-run cost visibility on everything.

The afterlife is the part I trust most. The weekly brief kept running after the engagement closed because the sponsor asked for it, more than once. The alumni movement cycle ran monthly for months after handover, and when the leader who commissioned it left the firm, he handed the monthly cycle to a successor on his way out. A system that outlives its own sponsor's tenure has stopped being a pilot. And one senior partner asked me, unprompted, whether I would run these signal subscriptions as a product across customers. That question is the whole expansion thesis, asked by the buyer.

The operating consequence is the part I care about: new requests extend the platform instead of starting over, and every new agent shipped faster than the one before it, because the shared work was already done.

Common questions

Could one assistant with a long prompt have done this?

It would have worked for the first team and failed at the second. One prompt cannot keep shared research behavior separate from one leader's preferences, cannot report cost per function, and cannot improve for one team without changing output for every team. The platform design is what let the build spread.

Who runs it day to day?

The firm's own people. Research operators check runs and correct audience configuration. Division managers set priorities and owners. Senior leaders read briefings and submit requests. That is the point of the handover: nine documented agents the firm operates itself.

How does a new intelligence request get added?

Through one intake: audience, question, sources, cadence, and owner, recorded before any work begins. That rule exists because the one tool built without it took six-plus revision cycles. Overlapping requests become platform additions instead of another standalone tool.

If your company has recurring questions that keep dying as one-off research requests, this is the shape of system that fixes it: focused agents, shared behavior, human judgment kept where it belongs, and a cost line on every run.