The AI Expertise Layer
Don't outsource all AI to foundation models — keep expertise and measurement to break the quality glass ceiling.
- Difficulty
- Advanced
- Time to result
- ~months to results
- Steps
- 4
- Confidence
- 88%
A framework for building AI products that avoids two extremes: the old 'hire data scientists for every small project' and the new 'LLMs will solve everything'. LLMs have huge utility but can't do specialized tasks like deal prediction, and without measurement you can't tell if V2 beat V1. Reshef prescribes retaining core AI expertise (data scientist, prompt engineer) and building measurement/eval rigor so quality can actually improve over time rather than stalling at a first-draft ceiling.
Origin
Eilon Reshef, whose Gong was built on machine learning before it was 'cool', articulates this from years of shipping AI products; he references Figma's naming of their AI feature 'first draft' as a useful conceptualization of LLM output.
Core principles
- 01LLMs are a great first draft, not a finished product — name and design around that
- 02Specialized tasks (deal prediction) need purpose-built models; LLMs cannot do everything
- 03Without measurement you have V1 then V2 with no way to know if you progressed
- 04Even if you outsource the core work, you must retain expertise to judge what's doable and how to approach it
How to run it
- 1
Retain a data-scientist role to guide feasibility
Keep a data scientist (full-time or advisor) to answer: is this an LLM task or a purpose-built model? What input is needed? How long will it take?
Pro tip Data scientists needn't all be in-house — advisors can cover the guidance function.
Watch out Don't assume the foundation-model company will handle your specialized problems — LLMs can't predict deals or other highly-specialized outputs.
- 2
Build measurement and a judgment/eval system
Install people who know how to measure whether one model or prompt is better than another — e.g. an Elo-style ranking (as used in chess) plus a way to put outputs in front of customers.
Pro tip Data scientists may not know if a given account brief is 'right', but they know what toolset and metrics you need to measure it.
Watch out Skip measurement and you hit a glass ceiling — you can't tell V2 from V1 and can't advance.
- 3
Add a prompt-engineering / optimization function
Keep someone (not necessarily full-time) who actually works with the LLMs — optimizing prompts, finding edge cases, ranking and improving outputs over time. This is a real technical skill.
Pro tip Gong's customers report its AI is more accurate than others — partly from models built/fine-tuned in-house, partly from prompt rigor and edge-case optimization.
- 4
Conceptualize output quality into the product design
Design the workflow and user expectations around whether the LLM is 90% accurate or 50% — the product looks different in each case. Frame it (like Figma's 'first draft') so users know what to expect.
Pro tip Embedding small AI-specialist teams inside pods lets them iterate quickly on LLMs vs SLMs as the field changes weekly.
In the wild
Gong has a deal-prediction model that LLMs cannot replicate because the task is highly specialized. Asking an LLM 'what does a good sales cycle look like' won't produce it — a purpose-built model does.
→ Retaining specialized modeling capability where LLMs fail.
Customers tell Gong its AI is more accurate than competitors'. Reshef attributes this to a combination of models built from scratch and fine-tuned (from in-house AI expertise) plus the rigor put into optimizing prompts and finding edge cases.
→ A measurable accuracy advantage over competitors.
Common mistakes
Assuming LLMs solve everything
Treating LLMs as a universal solution leads to spending many hours coaxing them into specialized tasks they fundamentally can't do, like deal prediction.
Shipping without measurement
Going straight to an LLM for V1 with no metrics means you can build V2 but have no way to know whether it's actually better — you stall at a quality ceiling.
Is it for you?
Best for
Product teams building AI features who are tempted to fully outsource intelligence to foundation models
Not ideal for
Truly generic, low-stakes LLM use cases where a first-draft output is genuinely good enough and no advancement is needed
From the transcript
“don't don't assume it does it does everything you still some need some of the core competencies of uh of AI so you do want…”
“we have a deal prediction model lmms cannnot predict deals because it's like very very specialized”
“if you don't have measurements like in the old machine learning whatever metrics you use you're not going to advance you're going to have V1…”
“figma calls their AI feature like first draft H which is a term I like because they kind of realize it's not best it's not…”
From the episode
Inside Gong: How teams work with design partners, their pod structure, autonomy, trust, and more
Eilon Reshef (co-founder and CPO)