Andrej Karpathy Feels Behind as a Programmer Due to AI's Rapid Advancement
He describes this new state as moving into 'vibe coding,' where AI assistance has progressed from suggesting code snippets to managing coherent, agentic workflows.
Andrej Karpathy discusses the revolutionary shift in AI's role, from code assistant to autonomous agent, redefining software development.
He describes this new state as moving into 'vibe coding,' where AI assistance has progressed from suggesting code snippets to managing coherent, agentic workflows.
Initial AI coding tools required significant manual editing for small code chunks, whereas by December, models began executing complex, multi-step tasks reliably, Andrej Karpathy notes.
By December, he observed that the latest models could generate complete, functional code chunks without fine-tuning, allowing him to assign complex, multi-step tasks that the AI reliably executed, leading him to invest significant time in new side projects.
He notes, "I just start to notice that with the latest models uh the chunks just came out fine. And then I kept asking for more, and just came out fine. And then I can't remember the last time I corrected it."
Software 2.0 previously introduced learned weights through data sets and neural networks, and now Software 3.0 leverages LLMs as interpreters controlled by context windows and prompting.
Karpathy states, "software 3.0 is kind of about uh you know, your programming now turns to prompting and what's in the context window is your lever over the interpreter that is the LLM."
Andrej Karpathy uses the OpenClaw installation as a case study to demonstrate how Software 3.0 approaches programming through instructions rather than fragile, platform-specific bash scripts.
He explains that the Software 3.0 approach involves providing an agent with high-level goals, allowing it to use its intelligence to observe the environment, install software, and debug autonomously without requiring precise, step-by-step instructions from the user.
Karpathy notes, "you don't have to precisely uh spell out, you know, all the individual details of that setup. The agent has its own intelligence that it packages up and then it kind of like follows the instructions."
By utilizing a 'menu-gen' project, Andrej Karpathy demonstrates that neural networks can now render intermediate application code unnecessary by using a single LLM to overlay content directly onto pixels.
The traditional method required building an entire app with separate components for optical character recognition (OCR) and image generation, creating an intermediate layer of code.
In contrast, the Software 3.0 approach utilizes a single Large Language Model (LLM) to directly overlay content onto pixels, demonstrating that neural networks now perform the heavy lifting and render intermediate app code unnecessary.
He explains, "The software 3.0 paradigm is a lot more kind of raw. It just um your neural network is doing more and more of the work, and your prompt or context is just the image, and the output is an image, and there's no need to have any of the app in between."
Andrej Karpathy contends that modern automation transcends traditional coding, as LLM-based knowledge bases enable computers to recompile and reframe unstructured data into entirely new products.
He highlights LLM-based knowledge bases as a new functionality that did not exist previously, enabling computers to recompile and reframe unstructured data, unlike the rigid, structured data of older code.
Karpathy emphasizes that focusing merely on speeding up existing processes overlooks the potential for entirely new products and capabilities that LLMs bring, stating, "It's not just even about code. So, previous code worked over a kind of like structured data, right? And uh you write code over structured data. But like for example with my LLM knowledge bases project... this is not something that could exist before."
The framework identifies roles as 'highly verifiable' or susceptible to automation, prompting a re-evaluation of those professional domains.
Andrej Karpathy explains that traditional computers automate tasks that can be precisely specified in code, while Large Language Models (LLMs) excel at automating what can be verified.
Frontier labs train these LLMs using reinforcement learning (RL) with verification rewards, which leads to 'jagged' capabilities, meaning models demonstrate peak performance in specific, highly verifiable domains such as mathematics and coding.
Karpathy elaborates, "basically like traditional computers can easily automate what you can specify in code. And uh, kind of this latest round of LLMs can easily automate what you can uh, verify in a certain in a certain sense."
Andrej Karpathy underscores that the 'jagged' performance of Large Language Models requires users to remain actively involved rather than blindly trusting outputs.
Because there is no comprehensive manual for these evolving systems, users should treat Large Language Models as tools that require continuous exploration of their strengths and weaknesses, Karpathy argues.
Karpathy states, "to whatever extent these models are remain jagged, it's an indication that number one, maybe something slightly off. Or number two, you need to actually be in the loop a little bit and you need to treat them as tools and you do have to kind of stay in touch with what they're doing."
Andrej Karpathy explains that verifiability makes a problem domain tractable within the current reinforcement learning (RL) paradigm for Large Language Models (LLMs), enabling founders to apply significant RL efforts.
Founders can develop their own fine-tuning processes to ensure the technology meets specific business needs, even if major labs do not prioritize that verifiable domain, Karpathy explains.
Karpathy states, "verifiability makes something tractable in the current paradigm because you can throw huge amount of RL at it. So, maybe one way to see it is that uh that remains true even if the labs are not focusing on it directly."
Andrej Karpathy believes that almost all tasks can ultimately be made verifiable to some extent, making them amenable to automation.
Creative tasks like writing could be verified by a council of Large Language Model judges, shifting the debate from automatability to the ease of implementation, according to his perspective.
Karpathy explicitly states, "I do think that ultimately almost everything can be made verifiable to some extent, some things easier than others."
Karpathy states, "I think this is magnified a lot more. 10x is not the speed up you gain. And I think it does seem to me like people who are very good at this peak a lot more than 10x from my perspective right now."
Andrej Karpathy emphasizes that humans must retain control over aesthetics, judgment, taste, and high-level oversight even as agents handle execution.
He explains that the inherent unpredictability of agents necessitates human intervention to manage disparate systems and requirements, ensuring overall coherence and quality.
Karpathy states, "You basically still have to be in charge of the aesthetics, the judgment, the taste, and a little bit of oversight."
He points out that relying on common identifiers like email addresses for database joins is prone to failure in agentic systems, emphasizing that developers must provide precise core specifications to prevent such logical errors.
Karpathy questions, "this is the kind of thing that these agents still will make mistakes about. It's like why would you use email addresses to try to cross-correlate the funds?"
Andrej Karpathy notes that AI models frequently produce code that is redundant, heavily reliant on copy-paste, and often fragile, which can be concerning for developers.
He attributes this to models struggling to simplify tasks because simplicity is not adequately rewarded in reinforcement learning, necessitating human developers to ensure code quality and maintainability.
Karpathy remarks, "sometimes I get a little bit of a heart attack because it's not like super amazing code necessarily all the time and it's very bloated."
He explains that AI's intelligence is shaped by data and reward functions, not by biological drives, highlighting their 'jagged' forms of intelligence.
This analogy helps users maintain realistic expectations for agent behavior, reminding them that AI operates differently from living organisms.
Karpathy states, "we're not building animals. We are summoning ghosts. Um and these are jagged forms of intelligence that are shaped by data and reward functions, but not by intrinsic motivation."
Andrej Karpathy argues that current software frameworks and documentation are fundamentally designed for human interaction, not for Large Language Models (LLMs), creating significant hurdles for agentic systems.
He views existing setup processes, such as DNS configuration, as particularly challenging for agents due to their human-centric design.
Karpathy advocates for designing future systems with sensors and actuators that are directly legible to agents, aiming for a future where agents can interact with infrastructure autonomously.
He comments, "Everything is still fundamentally written for humans and has to be moved around. I still use most of the time when I use different frameworks or libraries or things like that. They still have docs that are fundamentally written for humans."
He envisions a scenario where personal agents will interact with other agents to negotiate meeting details and logistical arrangements, reducing the need for direct human-to-human coordination.
Karpathy states, "I do think we're going towards a world where there's agent representation for people and for organizations, and you know, I'll have my agent talk to your agent to figure out some of the details of our meetings or things like that."
자막에서 근거가 되는 대목을 찾아 답합니다.
Sequoia Capital의 다음 글도 받아볼까요?
Sequoia Capital에 새 영상이 올라오면, 방금 읽으신 것처럼 정리해서 메일로 보내 드릴게요.