⚠ Synthetic pre-research — AI-generated directional signal. Not a substitute for real primary research. Validate findings with real respondents at Gather →
Projected from interview analyses using Bayesian scaling. Treat as directional estimates, not census measurements.
Side-by-side comparison of sentiment, intent, buying stage, and decision role across all personas.
Complete question-by-question responses with per-persona analysis. Click any respondent to expand.
Alex is a pragmatic, technically literate CTO navigating real procurement decisions around LLM adoption in a B2B SaaS context. His primary concerns are data residency guarantees, API and version stability, and the untrustworthiness of current benchmarks — all of which are actively blocking deeper commitment to frontier models. He is not opposed to these tools; he is using them selectively while routing sensitive workloads away from providers who cannot meet enterprise data requirements. His tone is measured and analytical, not enthusiastic or alarmed. He sees meaningful progress over the past 18 months but identifies specific, concrete gaps — particularly version drift and evaluation infrastructure — that remain unsolved. His unprompted flag about context window economics at scale suggests he feels the industry conversation is misaligned with the operational realities practitioners actually face.
Tell me what's top of mind for you on this topic right now — what are you wrestling with?
Right now the thing I keep coming back to is data residency and what actually happens to our prompts. We're a B2B SaaS company, so our engineers are naturally going to feed context into these tools — customer data, schema details, internal architecture decisions. And the moment that happens, I need to know exactly where it's going, whether it's being used for training, and what the contractual guarantees are. That's not hypothetical for us, it's an active procurement conversation. The second thing is model fragmentation. There are enough capable models now that picking one feels almost arbitrary, but the switching costs are real once you've built tooling around a specific API. So we're trying to make deliberate choices rather than just defaulting to whatever's most prominent, but the evaluation frameworks aren't mature enough to make that easy. Benchmarks don't map cleanly to our actual use cases.
What's the most important thing you need to understand or solve here?
For us, the core question is: which models can we actually build on top of without worrying that the foundation is going to shift underneath us? Capability matters, but consistency and reliability of the API are what I'm really evaluating. The secondary piece is data handling — we're a B2B SaaS company, so our customers' data is flowing through whatever we build, and I need to be confident about where that data goes and how it's used for training. That's a real conversation we have to have with every vendor before we commit to anything meaningful.
What does 'good' look like to you — and how far are you from that today?
For us, "good" means an LLM that we can deploy in our own environment, has predictable API behavior, doesn't silently change its outputs between versions, and gives us meaningful control over data residency. That's the bar. How far are we from that? Closer than we were 18 months ago, but not there. The frontier models — GPT-4, Claude — are genuinely capable, but we're routing sensitive workloads around them because we can't get on-prem deployment. The open-source side is improving fast, but the benchmark problem is real: there are too many models, the evals aren't trustworthy, and we end up spending engineering time just figuring out what to test against what. That's not a solved problem. The gap I feel most acutely right now is version stability. We'll tune a workflow around a specific model behavior and then something shifts in a provider update and we're chasing it. That's the kind of operational friction that doesn't show up in demos.
What would change your perspective on this entirely?
That's a fair question. Probably two things. First, if the data residency and sovereignty story got significantly cleaner — right now, even with enterprise agreements, there's enough ambiguity around where data actually lives and who can access it under what legal framework that I stay cautious. If a provider could give me something closer to the on-prem guarantees some of our larger customers demand, that changes the calculus a lot. Second, and this is maybe more fundamental — if benchmark quality actually improved to the point where I trusted them. Right now there are too many models, the benchmarks are all over the place, and I end up defaulting to the brand-name options not because I've validated them but because there's no reliable signal for anything else. If the evaluation infrastructure got more trustworthy, I'd be willing to explore the stack more aggressively rather than just defaulting to what's safe.
What question are you not being asked that you wish someone would ask?
The operational cost of context windows at scale. Everyone wants to talk about which model is smartest or which benchmark scored highest, but nobody's asking "what does it actually cost you when your engineers are feeding 50k-token contexts into production calls dozens of times a day?" That math gets ugly fast, and it changes which model you actually choose for a given use case — not because of capability, but because of economics and latency combined.
"We'll tune a workflow around a specific model behavior and then something shifts in a provider update and we're chasing it. That's the kind of operational friction that doesn't show up in demos."
Jordan is a Senior PM at a fintech startup wrestling primarily with model consistency degradation over time — not catastrophic failures, but gradual erosion of reliability that undermines engineer trust and slows adoption. The sentiment is measured and pragmatic: progress has been made versus a year ago, but a meaningful gap remains. The core concerns are twofold: (1) predictable, stable model behavior across updates, and (2) compliance-compatible data handling that determines which providers are usable at all. Jordan also surfaces a distinctive upstream insight — that engineer-level trust in models shapes which product features get built and how much human-review overhead gets baked in, a causal chain the industry is not yet asking about. Perspective-changers would be demonstrated consistency across version updates and rigorous domain-specific evals relevant to fintech work.
Tell me what's top of mind for you on this topic right now — what are you wrestling with?
Right now the thing I keep coming back to is consistency. We've had engineers on our team adopt a couple different LLMs — some are on Copilot for code, some are experimenting with Claude for longer reasoning tasks — and what I hear most is frustration when a model that was working well for a specific workflow just... stops performing the same way after an update. It's not a dramatic failure, it's more like a slow erosion of trust. And from a PM perspective, that's hard to manage because it's not a bug I can put in a ticket. The other piece I'm wrestling with is how to think about model selection at a team level versus individual preference. Right now it's pretty organic — people are gravitating toward what works for them — but I'm wondering if we need to be more intentional about it, especially as we think about data handling and what's actually going through these APIs.
What's the most important thing you need to understand or solve here?
For us, the core question is really about trust calibration — specifically, which models our engineers actually feel confident handing real work to versus which ones they treat more like a fancy autocomplete. We're a fintech startup, so correctness and consistency matter a lot. An LLM that does great work for two months and then starts degrading or behaving differently on the same tasks is a real problem for us — not just an annoyance. If engineers can't predict when a model is going to be reliable, they stop relying on it, and then we've lost the productivity gains we were counting on. The secondary piece is around data handling. We're dealing with financial data, and there's a real conversation happening about which providers we can actually use at all given our compliance posture. So trust isn't just "does it produce good output" — it's also "can we even send this data to that provider."
What does 'good' look like to you — and how far are you from that today?
For us, "good" means the engineers on my team are reaching for an LLM as a genuine accelerator — not fighting with it or working around its limitations. Concretely, that looks like: model outputs that hold up across a multi-week project without regressing on simple tasks it handled fine at the start, and responses grounded enough that engineers can trust them without spending equal time verifying. Where we are today? Closer than a year ago, but still a gap. The consistency piece is probably the biggest frustration I hear from our eng team — the model does something well early on, then a few iterations in it starts producing noticeably worse outputs on the same class of problem. That erosion of trust is real, and it slows adoption. So I'd say we're maybe 60-70% of the way there on the trust dimension, and the remaining gap is mostly reliability and predictability, not raw capability.
What would change your perspective on this entirely?
That's a fair question. I think if one of these models started demonstrating really consistent behavior over time — like, you upgrade to a new version and it doesn't suddenly regress on tasks that were working fine before — that would shift how much I trust it by default. Right now a lot of the skepticism I hear from our engineers comes from that inconsistency, where something works great for a few months and then quietly gets worse without explanation. The other thing that would move the needle for me is clearer data on how models perform on domain-specific tasks relevant to fintech — not just general benchmarks. If a provider published rigorous, reproducible evals on the kinds of structured reasoning and compliance-adjacent work our team actually does, that would be more convincing than any marketing claim.
What question are you not being asked that you wish someone would ask?
That's an interesting one to sit with. I think it's something like: "How does your engineering team's trust in a model actually shape which features you end up building?" Because from where I sit, if the engineers on my team don't trust a particular model's output for a given task — whether that's code generation, architectural suggestions, whatever — that skepticism filters all the way up into our roadmap. We'll deprioritize features that rely heavily on that model or we'll build in more human-review checkpoints, which has real cost and velocity implications. But nobody's really asking about that chain of influence. Most of the conversation stays at "which model scores best on benchmarks" rather than "how does ground-level trust actually propagate into product decisions."
"If engineers can't predict when a model is going to be reliable, they stop relying on it, and then we've lost the productivity gains we were counting on."
Chris is a measured, analytically-oriented demand gen leader wrestling with a real organizational alignment problem: his engineering team has formed strong, stress-tested opinions about which LLMs they trust, but those views aren't crossing over into marketing tool and messaging decisions. His primary concern is credibility risk with a technically skeptical ICP — positioning around an LLM his own engineers quietly distrust. On pipeline quality, he's at roughly 60% on attribution clarity and wants to tighten ICP fit rather than chase volume. He's open to shifting his perspective on LLM selection but wants practitioner-level, apples-to-apples evidence from real workflows — not benchmarks or vendor claims. His tone throughout is pragmatic and self-aware rather than enthusiastic or frustrated.
Tell me what's top of mind for you on this topic right now — what are you wrestling with?
Right now the thing I'm most actively wrestling with is how much our engineering team's tool preferences are bleeding into our go-to-market decisions. Like, they're already forming strong opinions about which LLMs they actually trust for real work — not just demos — and that's starting to influence what we can credibly say in our messaging and what proof points land with technical buyers. From a demand gen perspective, that matters a lot to me because our ICP skews toward VP of Eng and technical leads, and those folks are skeptical by default. They want to try before they trust, and they're definitely comparing outputs across models informally. So if I'm running content or campaigns that position us around a particular LLM and our own engineers quietly think that model is weaker for our use case, that's a credibility problem. I don't have a clean answer yet on which models engineers actually prefer and why. That's genuinely murky from where I sit.
What's the most important thing you need to understand or solve here?
For us, the core question is whether we can actually trust LLM outputs when we're using them in pipeline-facing work. Like, if I'm running an AI-assisted content workflow or using an LLM to help synthesize intent signals, I need to know the model isn't just hallucinating confident-sounding nonsense. The "which LLM do engineers trust" angle is interesting to me because engineers are usually the skeptics in the room — they'll actually stress-test the thing, not just take a vendor's word for it. So understanding what earns their trust tells me something useful about where the technology actually holds up versus where it's still pretty fragile.
What does 'good' look like to you — and how far are you from that today?
For us, "good" means pipeline with a clear, defensible story behind it. Not just volume — I want to be able to say here's the channel, here's the CAC, here's the conversion rate at each stage, and have that hold up when the CFO pokes at it. Right now we're maybe 60% of the way there. The tracking is decent, but attribution across touchpoints is still messier than I'd like, especially when deals involve multiple stakeholders touching five different things before they ever talk to sales. The other gap is on the quality side. We can generate lead volume, but the signal-to-fit ratio isn't where it needs to be — too many accounts entering the funnel that were never realistic buyers. I'd rather have a tighter ICP and fewer, better conversations than hit a number that looks good on a dashboard but doesn't close.
What would change your perspective on this entirely?
If I saw real, consistent evidence that one model materially outperformed others on structured reasoning tasks — not benchmark scores, but like, actual workflow outputs that engineers at companies similar to ours were willing to show side-by-side — that would shift things for me. Right now a lot of the "Model X is better" conversation feels like it's based on vibes or cherry-picked prompts. If there were more transparent, apples-to-apples comparisons from practitioners doing real work, I'd have a stronger basis to actually advocate for a specific tool internally rather than just letting our eng team default to whatever they're comfortable with.
What question are you not being asked that you wish someone would ask?
That's a good one. I'd say: "How are your engineering and product counterparts actually using AI tools in their day-to-day, and does that affect which LLMs you trust for marketing use cases?" Because there's this gap on our team where engineering has very strong opinions about Claude versus GPT versus whatever's newest — they've actually stress-tested these things. Meanwhile on the marketing side we're kind of just picking based on what's familiar or what our tools have baked in already. And I think those two conversations rarely cross-pollinate. If I understood what my engineering team trusts and why, that would probably inform my own tool choices more than any vendor comparison blog would.
"If I'm running content or campaigns that position us around a particular LLM and our own engineers quietly think that model is weaker for our use case, that's a credibility problem."
Marcus is a pragmatic, measured adopter of LLMs in marketing workflows. His primary constraint is not skepticism about AI capability but a structural compliance and data governance wall: enterprise customer data residency requirements make cloud-hosted models a non-starter for high-value use cases. His secondary concern is production reliability — specifically hallucination and output consistency at scale — which he frames as more operationally important than benchmark scores or demo performance. He reports cautious, partial adoption (content drafting, competitive research synthesis) and describes progress over 18 months but acknowledges a meaningful ceiling imposed by data handling constraints. He is notably even-handed: his skepticism is directed at vendor data practices and the gap between promises and terms of service, not at the technology itself. He would shift his position if verifiable private deployment options matched cloud model quality, and if real-task performance (not benchmarks) proved consistent over time.
Tell me what's top of mind for you on this topic right now — what are you wrestling with?
Right now the biggest thing is figuring out which models we can actually put customer data near. That's not a philosophical question for us — it's a compliance and procurement question. We've got deals with enterprise customers who have data residency requirements, and every time we want to spin up a new AI-assisted workflow, we're hitting that wall. The secondary thing is just consistency. From a marketing ops standpoint, I need outputs I can trust at scale — if I'm using an LLM to help draft campaign briefs or analyze intent signals, I need it to not hallucinate half the time and be coherent the other half. So the trust question for me is less "which model sounds smartest in a demo" and more "which model behaves predictably when you put it in production."
What's the most important thing you need to understand or solve here?
For us, the core question is really about reliability in production use cases. Not "which model scores best on some benchmark" — but which one actually does what we need it to do consistently, without us having to babysit the output or heavily prompt-engineer around its failures. We're using LLMs in a few marketing workflows right now — content drafting, some competitive research synthesis — and the trust question is less about raw capability and more about predictability. If I'm going to scale something into a team process, I need to know the model isn't going to randomly hallucinate a competitor's pricing or go off the rails on tone. That consistency piece is harder to evaluate than people admit.
What does 'good' look like to you — and how far are you from that today?
For us, "good" means an LLM that's reliable enough that I'm not second-guessing its outputs before I use them in something external-facing. Like if I'm using it to draft a positioning doc or synthesize competitive research, I need to trust that it's not confidently hallucinating a competitor's feature set. We're probably closer than we were 18 months ago, but still not there. The bigger gap for me right now is actually on the data side — we're cautious about what we pipe into cloud-based tools, so there's a real ceiling on how deeply we can integrate any of this into workflows that touch customer data or deal intelligence.
What would change your perspective on this entirely?
Probably two things. One is if a model consistently outperformed on tasks that actually matter to my work — not benchmark scores, but like, real campaign brief generation, market segmentation logic, competitive positioning drafts — and I could see that repeatedly, not just in a cherry-picked demo. The other is data handling. Right now there's a meaningful gap between what vendors promise about how your inputs are used and what's actually in the terms of service. If a provider gave enterprises genuinely verifiable on-prem or private deployment options with the same model quality, that would shift the conversation pretty quickly. A lot of the skepticism I have isn't about the technology itself — it's about what happens to the data I'm feeding into it.
What question are you not being asked that you wish someone would ask?
The deployment model question doesn't come up enough. Everyone focuses on which model scores highest on benchmarks, but almost nobody asks "where does this actually run, and who owns the data it touches?" For a lot of the companies we work with or sell into, that's the more fundamental constraint — cloud-hosted models are a non-starter for certain data types, full stop. The model quality conversation is almost secondary to that.
"The model quality conversation is almost secondary to that — cloud-hosted models are a non-starter for certain data types, full stop."
Synthetic pre-research uses AI personas grounded in real buyer archetypes and (where available) Gather's interview corpus. It produces directional signal — hypotheses worth testing — not statistically valid measurements.
Quantitative figures are projected from interview analyses using Bayesian scaling with a conservative ±49% margin of error. Treat as estimates, not census data.
Reflect internal response consistency, not statistical power. A 90% confidence score means high AI coherence across interviews — not that 90% of real buyers would agree.
Use this to build your screener, align on hypotheses, and brief stakeholders. Then run real AI-moderated interviews with Gather to validate findings against actual respondents.
Your synthetic study identified the key signals. Now validate them with 150+ real respondents across 4 audience types — recruited, interviewed, and analyzed by Gather in 48–72 hours.
"Which LLMs do engineers actually trust most — and why?"