Why converging models and identical consulting playbooks make your own customer data the last real advantage in life science marketing.
Shownotes
A paper in Nature found that models trained on model-generated data degrade one generation at a time, until the output collapses towards sameness. The same thing is happening to marketing teams, because everyone is drawing from the same public pool and buying the same four-workstream implementation playbook.
Who this is for: commercial and marketing leaders at life science tools and diagnostics companies who are being asked what AI changes about their competitive position.
Matt Wilkinson and Jasmine Gruia-Gray work through why intelligence has become cheap while truth and trust have become scarce. They cover the convergence problem across models and consultancies, why 80 to 90 percent of company data sits unstructured and unread, and what a question first approach to unlocking it looks like in practice. Matt also sets out where grounded synthetic customers fit, and Jasmine makes the case for the conference booth as the most under-recorded source of buyer language in the business.
Key idea: models and the consultants deploying them are converging on the same public information, so the customer knowledge your organisation already owns is the last durable advantage you have.
What you will learn
- Why model collapse and consulting convergence produce the same commercial outcome for your category
- Why having voice of customer research is different from having it turn up where decisions get made
- How to run a question first approach to unstructured data instead of a twelve month mapping exercise
- Where grounded synthetic customers earn their place across R and D, positioning and sales enablement
- How to read secondary market research reports without believing the absolute numbers
- Why a social media storm now gets regurgitated inside large language models long after it passes
- What a marketing or product leader can do in the next fortnight with the data already sitting in the CRM
Chapters
- [00:04] Welcome and end of summer catch-up
- [00:40] Model collapse: photocopying a photocopy
- [01:46] The untouched pool sitting next to the public one
- [03:10] Watermarking, the EU AI Act and the consulting convergence
- [05:41] Data silos as the enterprise's next big challenge
- [05:56] You have voice of customer, but does it turn up where it matters
- [07:37] Question first, and grounded synthetic customers
- [09:04] Where to start and which data to prioritise
- [11:34] The ten year old problem hiding in your messaging
- [12:08] Secondary market research reports and social listening
- [14:29] Social media storms now live inside the models
- [14:54] The conference booth conversations nobody captures
- [15:39] Recording, privacy and better trade show follow-up
- [17:31] What to do in the next couple of weeks
- [19:38] Truth and trust in one sentence
- [20:49] The Buyer in the Loop and close
Keywords
life science marketing, model collapse, unstructured data, voice of customer, grounded synthetic customers, AI convergence, EU AI Act, market research reports, trade show follow-up, product marketing, competitive moat, buyer insight
If this was useful, watch to the end, subscribe for new episodes of A Splice of Life Science Marketing, and read the full piece here. Matt's book, The Buyer in the Loop, goes deeper on keeping the buyer present through the whole commercial process.
Transcript
Matt Wilkinson and Jasmine Gruia-Gray discuss what happens when the models and the firms deploying them all start producing the same answers, and what a life science marketing team can do with the customer data it already owns. Speaker turns have been lightly cleaned for clear transcription errors and duplicated cross-talk between recording channels.
Welcome and end of summer catch-up
Speaker: Matt Wilkinson [00:04]
Hey Jasmine, how you doing?
Speaker: Jasmine Gruia-Gray [00:06]
I'm great, thanks Matt. It's the end of August already. How are you?
Speaker: Matt Wilkinson [00:13]
I'm good. I cannot believe July and August have disappeared already. I guess that's what happens when you self publish a book. Time just flies, and you have so much fun having so many different conversations that you wouldn't have had otherwise.
Speaker: Jasmine Gruia-Gray [00:28]
Well done.
Speaker: Matt Wilkinson [00:28]
Thank you. What I can't believe is that most of Europe's been on holiday, and all of that time I've been working hard. So there's obviously no rest for the wicked.
Model collapse: photocopying a photocopy
Speaker: Jasmine Gruia-Gray [00:40]
Apparently not. But I'm hoping you're going to reap lots of what you've been sowing over the summer. Speaking of which, let's get started on the most recent blog you wrote. I've done a little bit of research outside of this blog, and there's a paper in Nature that should worry anyone who thinks AI just keeps getting smarter on its own. Researchers found that when models get trained repeatedly on data other models generated, not real human data, just AI feeding on AI, the models start degrading. Details blur, errors compound, and the whole system starts collapsing towards sameness, one generation at a time. It's the AI equivalent of photocopying a photocopy of a photocopy until the page is barely legible.
The untouched pool sitting next to the public one
Speaker: Jasmine Gruia-Gray [01:46]
Now put that next to this situation. Most companies' data, somewhere between 80 and 90 percent of it, is unstructured. And most of them analyse less than a fifth of it, less than 20 percent. So you've got models trained on an increasingly shared, increasingly recycled pool of public information, converging toward each other. And sitting right next to that pool, almost completely untouched, is the one kind of data no model has ever seen and no competitor can scrape. Your own customers. That's the trade worth making. Stop feeding the machine the same public information everyone else is drinking from, and start feeding it the one thing that is actually yours and that you exclusively own. So Matt, that's basically the argument you made in your blog. Where did this whole topic start from for you?
Watermarking, the EU AI Act and the consulting convergence
Speaker: Matt Wilkinson [03:10]
Well, it's a couple of places. One, obviously you've got the watermarking conversation that's going on right now about the EU AI Act, and it's causing a bit of uproar. The AI companies love this, because it gives them a way to be able to differentiate. Hey, we already created this, we don't want to train on that data, or we want to downplay it. But realistically, the biggest challenge is that over the last couple of years, so since about 2023, the big four strategy houses have put more than 10 billion dollars into AI.
And there's a benchmarking firm that compared 10 of the big consulting companies side by side. And basically, underneath every single firm, they've got the same four work streams. They assess maturity, prioritise use cases, deploy on a governed platform, upskill the team. Now, if you're using multiple models, and I regularly am using ChatGPT and Claude, the convergence of those models and the harnesses around them is getting closer and closer. We can then see that the answers you get are getting closer. So they're converging around the same point. So if you've then got the people that are helping you implement AI converging as well, we're going to start creating that sea of sameness. It's the worst possible thing that can happen to a marketing team.
And so if we're starting to ask those questions of AI and we can go, brilliant, it can create us a marketing strategy, it can do A, B and C, well, what's left for us? And the big thing, and I think some of the companies like Oracle had already been talking about this, but really the place where most of our unique moat is, is what we know about the market, the environment in which we compete, and our customers.
And I think that set of data is massively underutilised in most organisations. I know from going and speaking to folks at Oracle, they're desperately trying to help organisations unlock the power of that data. But as you said in your introduction, so much of that data is unstructured. It's hard for organisations to really understand and tap into. And sometimes that's because the data's in different tools.
Data silos as the enterprise's next big challenge
Speaker: Matt Wilkinson [05:41]
It's in different data silos. And trying to tie all of that together in a way that actually allows you to make sense of it, I think, is going to be the enterprise's next big challenge.
You have voice of customer, but does it turn up where it matters
Speaker: Jasmine Gruia-Gray [05:56]
So there's a lot there to unpack. Let's start with the situation where most life science marketing teams already have some form of voice of customer. Maybe it's a year old. They have call recordings that maybe are more recent, from sales interactions or FAS interactions. Doesn't that mean they already have the thing you're calling scarce?
Speaker: Matt Wilkinson [06:36]
Well, they've already got the raw material. But what they really don't have is that data actually turning up anywhere that matters. So they don't have it in a readily available format. They still have to go find the recording. They have to go interrogate the CRM. They have to go look for the voice of customer deck that's buried on a hard drive somewhere. What they don't have is all of that data in a way that they can readily access and use.
And so that's one of the things that many organisations are desperately trying to unlock, is how do we actually map this together in a way that allows us to enable teams to ask better questions and answer them based on knowledge that only we have, and supplement what's available from the outside world. So it's really about enabling organisations to unlock the power of everything that they have collected already.
Question first, and grounded synthetic customers
Speaker: Jasmine Gruia-Gray [07:36]
So what's the best way to do that unlocking, especially if the information is unstructured?
Speaker: Matt Wilkinson [07:37]
Well, I think there's a number of ways of doing that. And I think you have to look at a question first approach. So what are the most important things? Now, anybody that has heard me talk or say anything recently will know that I get pretty passionate about creating grounded synthetic customers. And that's where we bring as much of that voice of customer and knowledge about the customer to life by creating representations of customer groups. And that can be decision making units in a big B2B sale. It can be across a whole range of different end markets. But really, what I'm trying to help organisations do is unlock the power of the data that they already hold, and make it accessible to their teams in the tools that they already use.
And so that's really what the difference is. You've really got to be able to look across the data in those data silos, make sense of it, and have a process where you make sense of that data to solve specific problems, and then turn that into a compelling way that people can easily access it.
Where to start and which data to prioritise
Speaker: Jasmine Gruia-Gray [09:03]
Okay, so if the organisation has a whole bunch of data, as you said earlier, in disparate areas, where's the best place to start? And what data should they use and should they not use?
Speaker: Matt Wilkinson [09:04]
Well, I think the first thing is, what question do you need to answer? So let's start from a strategic perspective, and what challenge are we looking to overcome? What would it look like if we had the ability, and let's take the example of a synthetic customer, to ask our customers questions at every step of this development process, from research and development all the way through to sales? If I've got the ability to interrogate my customer at will, what difference does that make? That might help me make better decisions in R and D. It might help me bring better positioning to market. It might help my sales team better understand the key objections that they're going to face and be able to role play some of those.
So it's really about understanding what challenge are we looking to answer, and then looking at what data do we have that, if we were going to spend time and time again answering that for everybody, we would go back to. And then we start looking at what do we need to collect in addition to that to answer these questions.
And what data can we use over time? So it's really, it's not about which data do we not use, it's about which data do we prioritise. And so you're looking at recency, obviously, but historical data can be incredibly important as well. So if you've got data over different timestamps, looking at where things have come from can help people better understand where things are now, and maybe where things are going to be going. And you can maybe start to see if somebody's not really going to buy an answer because they've seen the same problem for 10 years that's gone unaddressed. Maybe they're not going to believe that all of a sudden that problem's disappeared. So we can start seeing those trends in how we look to solve problems for people. And maybe that's the big thing. That this problem's existed for 10 years. We might not think it's a big problem, but actually that workaround has been causing problem after problem for a decade.
The ten year old problem hiding in your messaging
Speaker: Matt Wilkinson [11:34]
If we can resolve that, that might be something that we really, really want to prioritise in our messaging, because actually people will go, I didn't realise that was possible. And all of a sudden you're able to compete in a completely different way, because you've got that insight of, this has been a problem for a long time. Rather than, that's a small problem, it only takes three minutes every run. Well, three minutes every run, add it up. Days, weeks, months, decades. That's a big problem.
Secondary market research reports and social listening
Speaker: Jasmine Gruia-Gray [12:07]
So coming back to the beginning, around your suggestion to understand the market overall in addition to your voice of customer, how do things like secondary market research reports fit into this, where these reports might go into detail about competitors?
Speaker: Matt Wilkinson [12:08]
So those market research reports are often very interesting. I've been interviewed for a number of them in the past, and they take so-called expert opinion, and then they create an aggregate across them. So they look at, you look like you might know the market for X. Tell me about it. What's going to happen here? What's going to happen there? You get interviewed, they interview a number of people, they get people to make predictions for what's going to happen in terms of compound annual growth rate.
And then they look at those reports and they go, right, this is what we think's going to happen. We've interviewed 10, 20 experts maybe. Here's a report. And they'll look at some other signals as well. They'll look at what companies are saying in their annual reports, they'll look at what people that they consider to be experts say, and then they publish these reports. So they're useful as an aggregate, but I don't know how much I would ever believe the absolute numbers that are provided in them.
The other thing that I think is really interesting is using social listening and looking at what people are saying online about brands and markets and companies. There are very often a lot of people that will be talking about one thing or another, and online that makes a certain difference in those forums. But what's even more important is that the more weight a specific argument has, the more uproar there is about an event, the more the models will train on that. And actually that will then influence the results that the large language models will give you the next time, after the next training runs happen. So we really have to be very careful about the conversations that are happening online now.
Social media storms now live inside the models
Speaker: Matt Wilkinson [14:29]
And even more so than we ever did in the past. These days, a social media storm is no longer just a social media storm that will pass. A social media storm is a social media storm that will pass and then be regurgitated and fed back to potential customers time and time again, without you ever knowing, because it's in a large language model.
The conference booth conversations nobody captures
Speaker: Jasmine Gruia-Gray [14:54]
The other parts of the conversations that I wish I'd had the ability to capture were the conversations in the conference booth. The questions that you weren't prepared for, for one reason or another, and being able to capture those very, very rich questions, whether it's the conversation with a future customer and a product manager, or a salesperson and an existing customer. I think those would really be worth gold.
Recording, privacy and better trade show follow-up
Speaker: Matt Wilkinson [15:12]
Absolutely. There are challenges around what I'm going to suggest next in terms of privacy, and can you really record every conversation, or should you? We've seen a bit of a backlash against some of the Meta glasses that do that sort of recording everything. But there are tools such as Plaud, P-L-A-U-D, and others that are AI transcription note takers. They record what you're listening to and then they transcribe it when you say, yes, I'd like to transcribe that. And they're fantastic at conferences where you might want to record all of the talks. Absolutely amazing experience for being able to record something, get the transcripts of what was said in the talk, and then bring that back.
You could imagine the same thing happening on a booth very, very easily. And then you could look at being able to use that data in two ways. The first, I think, is a huge place where you and I have spoken about trade shows a number of times in the past, where we could use that conversation data to actually deliver a more relevant follow-up than ever before. So rather than, hey, we spoke about A, B and C, would you like to schedule a call? Hey, we spoke about this stuff. Here's the exact answer to your challenges. These are some potential next steps. Able to take a much better approach to your follow-up.
And the other, of course, is that you can bring all of that back. You can use it to better prepare the teams on the booth anyway, to answer the questions that they weren't expecting. And then of course you can use that data to also look at, well, what are the questions that we need to be answering on our website, in our materials, in our outreach? All of those sorts of conversations, that data becomes just a gold mine for almost any marketing campaign you're ever going to run.
What to do in the next couple of weeks
Speaker: Jasmine Gruia-Gray [17:31]
Yeah, fully agree. This is where I plead with product marketers and marketing managers. Please, please go to conferences, stand in your booth, do listening, engage in conversations, because you'd be surprised at the wealth of information that you can come back to the office with. So with that, Matt, for someone listening who runs marketing or is in product management, what do they actually do in the next couple of weeks with all the unstructured information they have?
Speaker: Matt Wilkinson [18:18]
Well, I'd start looking at trying to pull that together. First off, looking at what unstructured data do I actually have? That might be interview transcripts, it might be the calls that get logged on a CRM. Maybe they are transcribed, and maybe those transcriptions are in the CRM itself. Maybe there are support tickets that you can see that are looking at the questions that customer services get asked. There might actually be knowledge bases that are automatically being created for those. And look at all of those and start thinking, how can I use this to better support the teams that marketing supports, to provide them information?
One form of that might be a synthetic customer, but it might just be creating a well-grounded database that you use, or a knowledge base, I should say, that is then shared as part of marketing's deliverable internally to the other departments. So that data is no longer just sat in disparate silos, but is all of a sudden an asset that teams can use.
Truth and trust in one sentence
Speaker: Jasmine Gruia-Gray [19:26]
Yeah, well said. So in one sentence, how would you summarise your blog post?
Speaker: Matt Wilkinson [19:38]
Great question. I think that we're moving into a space where the two things that are most important for organisations to chase after are the truths about our customers, and the trust that we can gain with our customers, by actually turning what we already know into something that enables us to serve our customers better.
Speaker: Jasmine Gruia-Gray [20:08]
Yep. Dig into your own customer data. That would be my one sentence. So the full piece is called Intelligence is Cheap, Truth and Trust Are Scarce. It's on the strivenn.com/thinking web page. And I really recommend reading the full blog post. It really is insightful and gives you a different perspective on AI.
The Buyer in the Loop and close
Speaker: Matt Wilkinson [20:49]
Thank you. And if you want to go deeper on how buyers disappear between the brief and the market, and how grounded synthetic customers might actually help keep customers in the room instead of rusting away on a hard drive, that's the whole subject of my book, The Buyer in the Loop. The link will be in the show notes, or you can find it on Amazon wherever you are.
Speaker: Jasmine Gruia-Gray [21:09]
Superb. Thanks for another great conversation, and thanks to everybody for listening to another episode of A Splice of Life Science Marketing.
Speaker: Matt Wilkinson [21:09]
Thank you. Bye for now.
Speaker: Jasmine Gruia-Gray [21:20]
Bye for now.
Q&A
We already have a voice of customer deck from last year. What do we actually do with it next week?
Pull it off the hard drive and give one person half a day to strip it down to the twenty questions customers actually asked, in their words. Put that list in a shared document your product and sales colleagues can open without asking permission. It costs nothing beyond the half day, and it turns dormant research into something a colleague can use on Monday morning. The payoff shows up in the next brief.
Everything we have is unstructured. How do we decide what to prioritise?
Start from one question you get asked repeatedly, such as why deals stall at evaluation. Work backwards to the data that would answer it: call recordings from the last two quarters, support tickets tagged at that stage, and win loss notes. Ignore everything else for now. Give it a week and one person. A narrow pass on a single question teaches you more than a twelve month data mapping exercise.
How is a grounded synthetic customer different from the persona document we already have?
A grounded synthetic customer is built from your own transcripts, tickets and call notes, so it answers in the language your buyers actually used. Most persona documents sit in a folder and get read once. To test it cheaply, take ten interview transcripts, load them into whatever model your company already licenses, and ask it to answer three objections your sales team hears weekly. Check the answers against a real customer call before trusting them.
We go to four conferences a year. How do we capture booth conversations without a privacy problem?
Ask before you record, every time, and say what the recording is for. Use a transcription device such as Plaud for talks and sessions, where consent is simpler, and stay with written notes on the booth until legal has signed off a process. Give one product marketer responsibility for logging the three questions nobody on the stand could answer. That list alone will rewrite your next FAQ page.
If models keep converging, what should we stop doing?
Stop briefing agencies and consultants from public market reports alone. Those reports feed the same aggregate everyone in your category is reading, and the growth numbers rest on a small pool of expert interviews. Add one section to every brief that quotes your own customers verbatim, sourced from a call in the last ninety days. One person can maintain it in an hour a month, and it changes what gets approved.