prompt engineers are yappers
karpathy calls it context engineering
👋 Hey, welcome to this week’s edition.
Quick poll before we get into it:
In 2023, Anthropic listed a prompt engineer role paying $175K to $335K and by 2025, Fortune called the job already obsolete.
The gold rush was built on the idea that the skill was compression i.e. find the right words and use as few tokens as possible.
That made sense when typing was the only way in.
Models also got much better in that time.
In 2023 you had to hand-hold by telling it to think step by step, give it a persona, and spell out the format.
If putting 4 paragraphs of background into a prompt took 8 minutes of typing, you’d cut the background.
Everyone did, and models spent 2 years getting half the picture.
Karpathy renamed the discipline to “context engineering” last June, and everyone moved on but I think the rename missed the bigger change, which is what I want to talk about today.
Why the yappers are winning
Speaking 400 words takes about 2.5 minutes while typing them takes closer to 8.
Once that gap opened up through dictation tools, people stopped rationing their background, and models started getting prompts with enough detail to work with.
Karpathy posted about his own version of this earlier this week.
If you do the math, a 10-minute ramble is about 1,500 words, so roughly 2,000 tokens.
Frontier models sit at 1 million token context windows, so your ramble is only 0.2% of the space available.
I’ve been doing something similar for months, and my outputs improved without me having to change models too much.
The only variable was how much of the problem I was willing to put into the prompt, and voice made that number go way up.
Before you throw out your keyboard
Jerry Liu, replied to Karpathy’s tweet saying he used to ramble too, then started feeling like he was “getting progressively dumber at writing.”
His fix was hand-typing distilled bullet points after reading the model’s output.
A study of 919 programming students that came out this month speaks to both sides of this.
Students who typed their prompts outperformed students who submitted raw, unedited voice prompts, but students who spoke and then edited their transcript before sending matched the typers.
The model is not telling you to talk more, but to speak the context out, read it back, clean up the tangents, then send.
Talk > see > edit > send.
This is relevant because a lot of voice tools auto-submit when you stop speaking.
Claude Code’s /voice streams a live transcript into the input box and waits for you to hit enter, while ChatGPT’s Advanced Voice Mode auto-sends on a pause and people keep losing their train of thought mid-prompt.
One preserves the edit step and the other throws it away, and the edit step is where the difference between coasting on the model’s thinking and doing your own shows up.
The setup worth trying
Context that would take 15 minutes to type comes out in 3 minutes of talking + 2 to clean up.
But the model gets a longer, more honest, and specific version of it than what most of us were putting in a year ago.
My workflow with Wispr Flow went this direction.
I hold the hotkey, talk through what I’m trying to figure out, watch the transcript appear, cut the noise, send.
The brief the model gets is richer than anything I’d have typed, and my outputs improved.
If you want to try the same loop, TPFXWISPR gets you 90 days free.
Upcoming Events
Founder of Founders: A Binny Bansal Exclusive
August 1 | Bengaluru
Register Here
Bits n Atoms: Build with OpenAI team
5th August | Bengaluru
Register Here
Basecamp | 6-12 August, Bengaluru
speakeasy w/ wispr flow
8th August | Register Here
Ship it | Replit x TPF
9th August | Register Here
Cafe Cursor
9th August | Register Here
Women in AI: In Conversation
10th August | Register Here
Feature Showcase Party by ElevenLabs
11th August | Register Here
Exclusive Jobs of the Week
Senior Product Manager
Lead Product Manager
Director Product
These and other roles open across top companies like Meesho, Google, Zomato & many more.
Bottom line
Talk > see > edit > send.
Talking is cheap, seeing is where you decide whether you’re offloading or doing your own thinking, and the edit is what makes the output yours.
Reply and tell me if you’re a yapper too. I have a feeling more people than expected are going to say yes.
Cheers,
Suhas 👋🏻
P.S. This gets better when the right people are in the room. Share it with one.





