L&D and enablement teams are racing to build and improve AI learning and coaching agents: tailoring them to specific use cases, enriching them with internal data, and connecting them to enterprise systems via APIs. But these investments can still feel like bets.
The question is: how do we know if efforts to improve AI agents actually pay off?
One of the bets that we are closest to at Primerli is content ingestion: feeding data and information to agents to give them more context. We are already working with clients to do exactly that: integrate Primerli’s industry-specific learning content into their enterprise AI agents.
To test the impact of that content on agent performance, we ran a head-to-head experiment. We compared generic, off-the-shelf coaching agents with agents enriched with Primerli’s industry content. This was an all-AI experiment: AI agents, AI personas as users, and AI models as judges.
The results were promising: Primerli-enhanced agents outperformed generic agents by 15–20% across our 4-part, 100-point rubric.

Primerli’s content improved agent performance on all four dimensions, but the biggest lift came from accuracy and insightfulness – the Primerli agent outperformed the generic agent by more than 20% on both dimensions.

We found this result intuitive. Primerli content is expert-vetted and therefore reliably more accurate. Plus, because the content is developed by industry experts and former management consultants, it emphasizes the points that are most insightful for client-servicing professionals. It does more than define terms; it focuses on the implications, trade-offs, and business impacts that leaders care about. This additional context set the Primerli-enhanced agents apart.
Examples of better agent performance
The Primerli content improved the agents' accuracy and insightfulness by 20%. To illustrate what exactly accounts for that, we’ve pulled a few short examples. These are drawn from the conversations we ran with the generic and Primerli-enhanced agents.
The context for each conversation was the same: a salesperson preparing for a meeting to discuss insurance claims and automation opportunities.
An example of more accurate: the buyer's true focus
- The generic agent presented a common but problematic mischaracterization of insurance claims, saying the goal is to “minimize what the company pays out…legally.”
- An industry insider would disagree, describing this as an unfortunate misunderstanding they often encounter. Instead, as the Primerli-enhanced agent correctly says, “there is nuance: they don’t just want to pay less – they want to pay the right amount, but as fast as possible.” If you are working with a claims leader, emphasizing ‘minimizing payouts’ would miss the mark. Fast and fair is their focus.
An example of more insightful: the tradeoffs of automation
- The generic agent described the benefits of automation in claims in simple, cost-reduction terms. Specifically, pointing to the opportunity to reduce overheads in processing claims, which are called LAE (loss adjustment expenses). The generic agent said: “Connect your workflow automation to LAE per claim…you’re not just making them faster, you’re making them cheaper to handle.” This is not wrong, but it is missing a big so-what.
- The Primerli-enhanced agent adds the missing nuance: “it gets counterintuitive: if you cut LAE too aggressively…you end up increasing indemnity, which is the actual payout to customers.” This trade-off is extremely important in the claims context, since the reason to have manual oversight rather than fully automate claims is that you might end up paying more.
For our users, who are professionals at leading tech companies and consulting firms, these gaps in accuracy and insightfulness are critical. They cannot settle for ‘close enough.’ The costs of misunderstanding a concept or having the wrong conclusion about what it means are simply too high. They can lose credibility instantly, damaging or ending working relationships.
The results were consistent across dozens of runs
Overall, the Primerli-enhanced agent outperformed the generic agent by 18%. This was true over dozens of runs, with different combinations of agent types, user personas, and LLM-based ‘judges.’ The performance impact of Primerli content was remarkably consistent, as you can clearly see in the score distributions below.
The Primerli-enhanced agent received higher scores across the board, with an average total score of 85/100. The Generic agent’s scores are lower across the board, with an average of 72/100.

Where do we go from here?
The truth is that generic agents are already very impressive, and their performance is getting better every day.
But out-of-the-box agent performance is absolutely not the gold standard. In fact, simply using generic agents leaves a lot on the table, especially for enterprise use cases where accuracy, nuance, and business context matter.
The signal from this experiment was clear: high-quality industry learning content improved agent performance. Of course, this was only one experiment. There is much more to test across content sources, industries, agent designs, user personas, workflows, and evaluation methods.
For us, the exciting part is not just the 15–20% lift. We’re also excited that the path to measuring agent improvement is becoming clearer.
That matters because without an approach to measurement, L&D and enablement teams are left making bets: adding more data, features, and integrations without knowing whether those changes actually improve performance.
Having a clear methodology and approach will help L&D and enablement teams work smarter, not harder: defining what good looks like, establishing a rubric, testing specific changes, and measuring whether those changes improve the user experience.
The next frontier is not just building AI agents. It is learning how to improve them deliberately.
If you are interested in Primerli’s content or our AI agent integration work, we would love to talk.
Read more about what we did
The approach
We wanted to answer two questions: does industry context improve AI agents? And, does that improvement vary for different types of agents?
Our customers are asking themselves these same questions, albeit a bit more broadly: what types of agents to build, and how exactly to improve them?
So we ran this experiment with two types of agents, reflecting the most common use cases we are hearing from customers:
- The Industry Tutor: Instructed to build industry fluency by teaching the basics of the industry in a tailored way
- The Deal Prep Coach: Instructed to help the seller prepare for a specific meeting by surfacing the use cases, buyer priorities, and business implications most relevant to the client
We ran a head-to-head test between different versions of both AI Agents: versions without any Primerli content as context (the Generic Agent) and versions enriched with Primerli content (the Primerli-enhanced Agent).
This was an all-AI experiment: the Agents are AI, the ‘test’ user personas are AI, and the judges are AI. All were prompted with detailed guidelines and instructions. All agents and user personas had the same context: an upcoming meeting with the Chief Claims Officer (CCO) of an Insurance Company.
To make the experiment more robust, we wanted to test whether the results held across different types of users and repeated runs of the same basic conversation. A single interaction can be noisy: one persona might ask better follow-up questions, or one run might produce a stronger or weaker response by chance. So we created four distinct sales personas and ran each conversation three separate times. Each conversation was then scored independently by two different LLM judges using the same rubric and evaluation standards.

The results
The results were remarkably clear and consistent. The Primerli-enhanced agents outperformed the generic versions across the board. This was true for both types of agents, all user personas, and according to both AI judges.
Adding Primerli content improved agent performance by 15-20%+
The impact was strong in both types of agents. Primerli content improved Industry Tutor performance by 21% and Deal Coach performance by 15%. The fact that our industry content had an even bigger impact on an agent designed to teach the industry makes intuitive sense. Primerli content is particularly well-suited for learning and teaching moments, since our primers are built by experts to teach the basics as effectively as possible.

The results were consistent across both LLM judges and all four user personas, with no meaningful variation by judge or persona.
Frequently Asked Questions

.png)
.png)
.png)

