Introducing the Gemini Interactions API and Managed Agents
The primary goal of these new tools is to address the widening gap between the evolving capabilities of AI agents and the constraints imposed by legacy API structures.
Google DeepMind introduces new tools to streamline AI agent development, addressing challenges with evolving AI capabilities and legacy API structures.
The primary goal of these new tools is to address the widening gap between the evolving capabilities of AI agents and the constraints imposed by legacy API structures.
Early model interaction primarily involved simple, single-turn message-response cycles, which were straightforward but limited in scope.
The introduction of function calling allowed models to produce predictable JSON structures, enabling more sophisticated tasks such as data extraction and specific operations.
Modern AI agents have evolved to reason over extended periods, utilizing multiple tools, critically reflecting on their output, and interacting dynamically with their environment, moving beyond specialized tools to broad execution environments like bash.
The speaker demonstrated AI Studio's efficiency by creating a mobile-responsive website for his presentation notes in just two minutes.
AI Studio provides a simplified user interface for parameter tuning and testing models like Gemini 3.5 Flash, with Google offering a generous free tier for experimentation, including advanced features such as speech-to-speech translation.
I actually like to point out that my speaker notes today were actually coded on AI Studio. I vived it out 2 minutes before this talk.
Google's current AI models span a wide range, from lightweight units like Gemini 1.5 Flash to highly complex research agents, presenting a challenge for unified management.
The existence of diverse API endpoints and deeply nested data structures has previously hindered first-party integrations, motivating the development of the Interactions API to unify Gemini's utility across the entire Google ecosystem.
Previously, manual management of 'thought signatures' was fragile, with even minor errors like an extra whitespace invalidating the model's cache.
The new API replaces this manual process with a single interaction ID, which automatically preserves context across multiple turns, a critical feature for maintaining consistent performance in the latest Gemini series models.
A demonstration illustrated the API's capability to generate variations of a person in different global locations from a single photo.
The Interactions API allows for spawning multiple agentic calls from an initial request and enables seamless transitions between models—such as from Gemini 1.5 Flash to the new Omni model—while maintaining shared context and images.
It's making it possible to build more complex, rich, and just these really complex agentic applications.
Legacy API endpoints often required navigating complex, deeply nested objects to retrieve outputs.
The new system uses an 'output.type' discriminator to clearly mark return data, simplifying parsing and allowing developers to switch modalities, such as from audio to image generation, by simply adjusting the response modality and generation configuration.
The new API enables developers to mix and match both built-in and custom tools, enhancing the flexibility of agent capabilities.
An example showcased an agent utilizing Google Search to retrieve the latest security reports for a React application, demonstrating the integration of external tools.
The model can use a context tool to retrieve information from websites via URL, extending its capabilities beyond static training data.
Custom tools, such as a 'file incident' function, allow for autonomous reasoning within a single API call, enabling the model to determine necessary information and generate responses independently.
It replaces the legacy output array with a tight discriminator, clarifying model outputs and function calls at each step and explicitly tracking content types like audio and video to facilitate multimodal pipelines.
Antigravity serves as a standardized harness utilized across Google's product suite to streamline AI agent deployment and management.
Managed Agents provide a persistent remote sandbox, eliminating the need for developers to manage infrastructure manually, allowing them to focus on tasks like autonomous repository analysis.
The system ensures state persistence by using environment IDs and interaction IDs, maintaining context across various API calls within the sandbox.
With support for loading sources from Google Cloud Storage buckets, GitHub repositories, and inline files, the platform ensures the model retains installed packages and files across API calls.
Persistence allows the model to retain installed packages and created files across subsequent API calls, enabling high-complexity tasks such as analyzing massive codebases with millions of tokens.
Developers can package tuned skills into a single folder and upload them, eliminating the need to manage underlying infrastructure or fine-tune the harness after local testing is complete.
This system dynamically injects tokens, such as GitHub tokens, into headers at runtime, preventing the model from ever directly accessing raw credentials and mitigating risks from prompt injection.
Agents can be created by freezing a specific set of files and network sources, allowing for highly customized configurations that can be iteratively refined through chat.
Users can manage up to a thousand named agents without incurring additional storage or sandbox costs, as payment is solely based on model usage, enabling flexible and scalable agent deployment.
Google DeepMind also provides a new 'interactions API' skill to facilitate the migration of existing coding agents to the new API structure, with ongoing maintenance and evaluation ensuring compatibility with evolving model versions like Gemini 2.0 and 2.5 Flash.
Answers come from the transcript, with the exact spot cited.
Want the next article from AI Engineer?
When AI Engineer publishes, we'll write it up like the one you just read and email it to you.
AI Engineer published 29 in the last 7 days.
Nebius Optimizes Open LLMs for ProductionAI Engineer25 minutes ago · 20:32 · 26 views · Created 14 minutes ago
Crusoe's Self-Healing AI TrainingAI Engineer1 hour ago · 16:54 · 1 views · Created 1 hour ago
Wearable AI Agent Built on Raspberry PiAI Engineer6 hours ago · 20:10 · 16 views · Created 6 hours ago