DevelopmentNew Release5 min readPublished September 26, 2026

Local inference changes where work runs, not who owns it

Google's Antigravity SDK Now Runs Local Models Offline

Google's Antigravity SDK can now run agents on local models with no internet connection. What changed, how to set it up and the limits Google states.

DA
Digital Applied Team
Research and practical guidance
CoverageThrough September 26, 2026

Google’s Antigravity SDK can now use local AI models for agent workflows. The September 23 announcement gives developers a way to keep inference on their own machine, including when the network is unavailable. That creates a useful option for code and documents that should stay local, provided the tools around the model follow the same boundary.

Editorial note: Prepared October 1 from information published by September 26, 2026. Later product developments are outside this article’s scope.

Key takeaways
  1. 01
    Choose the data boundary firstDecide whether task descriptions, filenames or results may leave the machine.
  2. 02
    Prepare before disconnectingDownload the runtime, model and task dependencies before testing offline operation.
  3. 03
    Measure accepted workHardware capacity, task quality and review time matter alongside avoided API calls.

01 — The evidenceWhat the announcement adds

Google’s launch post highlights Gemma 4 26B A4B through LiteRT and recommends more than 24 GB of VRAM or unified memory for that setup. It also identifies OpenAI-compatible local servers through LocalOpenAIAgentConfig. This is SDK support: it does not mean every Antigravity editor session has switched to local inference.

An SDK is a library used to build an application. The model produces responses, while the surrounding agent code decides how to call tools and continue the task. Moving the model changes where those responses are generated. It does not automatically decide which files a tool can read or where a tool can send results. For the wider product context, see our Antigravity desktop overview.

02 — Practical implicationsChoose fully local or hybrid deliberately

Fully local
Keep the task on the workstation
No cloud planner

Use this design when even filenames and task descriptions must stay on the device. Validate the complete workflow while disconnected.

Strict data boundary
Hybrid
Send a limited brief to the cloud
Local workers

Use this only when the information sent to the planner is explicitly approved. Review what workers return as carefully as what they receive.

Defined disclosure

Google demonstrates a hybrid workflow in which a cloud planner receives filenames and task descriptions while local workers handle source code. That is a vendor demonstration, not proof that an arbitrary hybrid application keeps every sensitive detail local. We have not run an independent performance or privacy benchmark of it.

The design question is concrete: which exact fields cross the boundary? A filename can expose a customer name. An error message can include a path, a query or a fragment of a document. A “summary only” return value can still disclose the content it summarizes. Review the outbound payload rather than treating the absence of a full source file as sufficient evidence.

For a hybrid pilot, use synthetic filenames and a disposable repository. Capture the requests sent to the cloud planner and inspect them before using real work. If the organization cannot authorize even that limited transfer, choose the fully local path instead of trying to phrase a prompt that promises privacy.

03 — Practical implicationsPrepare a reproducible local environment

The official Python SDK repository is the implementation reference linked by Google’s launch post. Record the package version, model artifact and runtime used for your pilot. A later package install can change behavior, so “we used the SDK” is not enough detail to reproduce the result.

Digital Applied preparation checklist; no local benchmark was performed for this article.
StepWhat to establish
InstallUse an isolated environment and record the installed packages.
ModelDownload the intended artifact and record its identity and location.
WorkspaceCreate a disposable directory containing only approved inputs.
ToolsAllow only the operations needed for the task.
Offline testDisconnect and repeat the task with its dependencies already present.

Use Google’s dated setup instructions for the exact installation and configuration syntax. The LiteRT route uses LiteRTAgentConfig; the compatible-server route uses LocalOpenAIAgentConfig. Check that the endpoint you configure is actually local. A localhost address inside a container refers to that container, which may not be where the model server runs.

Start with reading a small test directory and producing a summary. Add file writing only when that simple loop behaves as expected. Keep the first write task recoverable: a generated report or a patch in a throwaway checkout. The purpose is to establish that configuration, inference and tool execution work together before evaluating a larger task.

04 — Practical implicationsLocal execution still needs permission limits

A local model can operate on valuable files if its tools can reach them. Our sandbox and filesystem boundary guide explains why a working directory alone should not be treated as an access-control guarantee. Give the runtime only the files and credentials the task needs.

Review the demonstration\u2019s permissions

Google’s resource-monitor example includes an allow-all policy. Treat that as demonstration configuration. Choose and test permissions for your own task before connecting the agent to a real workspace.

Keep secrets out of the test environment. Review mounted directories, environment variables and inherited credentials, then check whether a permitted command can invoke another program with broader access. For a truly offline requirement, verify network behavior at the runtime or operating-system boundary. A promise in the prompt is not a network policy.

05 — Practical implicationsDecide from the work your machine finishes

Compare a fixed set of tasks with your existing process. For each, record whether the result passed, how long it took, peak memory use and the time a reviewer spent fixing it. Add the cost of hardware and operations to the decision, even though local inference avoids a per-request cloud API charge. An idle machine you already own and a new dedicated workstation have different economics.

Start with a task whose correctness is observable, such as making a small change that must pass an existing test. Avoid ranking the setup by how quickly it begins producing text. The useful result is the accepted patch or report, including any retries. Our model-routing guide covers that workload-based comparison; our AI transformation service helps teams design the evaluation.

Next step

Test one local workflow from input to accepted output

The new option is valuable when it fits the data boundary and the hardware. Prepare the dependencies, restrict the tools and run a task you can verify. Expand only after that complete workflow succeeds under the network conditions you actually require.

Agentic AI implementation

Build a workflow you can evaluate and control

Digital Applied helps teams connect AI capabilities to useful work, clear acceptance checks and responsible operating limits.

Task evaluationsCost visibilityControlled access
Start with one task

Define the pilot

  • →Approved source material
  • →A named reviewer
  • →A clear acceptance check
  • →Spending and permission limits
Questions and answers

Practical questions

Only if the tools and dependencies also work without network access. Test the whole workflow while disconnected.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source