Switching a coding model often means switching the tool around it: a different terminal interface, another configuration file, and another round of teaching the agent how your repository works. Together AI launched Together Link on October 5 with a more useful proposition: keep the coding agent you already use, but connect it to models hosted by Together AI.
The supported tools include Claude Code, Codex, OpenCode, Pi Code, Claude Desktop, and ChatGPT Desktop. Together’s launch promises more than 50% lower spend, but that is a vendor claim, not an independently measured result for your backlog. The interesting question is whether you can change the inference backend without giving up the workflow—and whether the cheaper backend still finishes the work.
Keep the Harness, Change the Backend#
A coding agent is more than a model. Its harness manages repository access, tool calls, permissions, conversation state, and the loop that turns a request into a change. Together Link separates that interface from the inference service behind it.
According to the current integration documentation, Link configures supported tools to connect to a hosted gateway; there is no local proxy or daemon. Terminal agents receive temporary, per-launch configuration, while desktop applications use separate Together Link profiles. That distinction matters: a reversible desktop integration still leaves a persistent profile, rather than disappearing when a terminal session ends.
This is hosted inference, not local model execution. Open weights do not mean your repository context stays on your laptop, and using Claude Code’s interface does not mean Claude produced the answer. My earlier Kimi K3 open-weight analysis covered the model side of that choice; Link addresses how to connect an existing workflow to such models.
Start With One Reversible Terminal Session#
The setup guide requires macOS or Linux, a Together API key, and an already-installed supported agent. Link does not install the coding agent for you. Its installer can install Bun if needed and places its commands in ~/.local/bin.
Together documents a pipe-to-shell installation command. I would download and inspect that script before executing it instead:
curl -fsSL https://link.together.ai/install -o together-link-install.sh
less together-link-install.sh
# Run only after reviewing the installer:
bash together-link-install.shThe installer can open the interactive launcher when it finishes. For subsequent setup and a fixed-model trial, the documented commands are:
togetherlink configure
togetherlink models
togetherlink --main zai-org/GLM-5.3 claudeconfigure saves and validates your Together key. Do not assume a key sitting in your project’s .env is enough: the docs explicitly say Link does not load those files. Keep the credential out of the repository and out of bug reports.
The flag order is important. --main belongs before the tool name, and the short aliases such as tclaude cannot be used to select a model this way. Passing an agent’s --model flag after the tool name is not an equivalent gateway selection; Link warns that it can be refused or ignored.
Pinning the model makes the first test easier to interpret. Start in a disposable worktree with a small task and your normal approval rules, inspect the resulting diff, and launch the original agent directly when you want to return to the original setup. I have not run a live Together Link benchmark for this article; these are source-checked setup instructions and a proposed evaluation method.
Auto Routing Has a Documentation Caveat#
The launch post says the router reads the first task and chooses a model once per session, preserving prompt-cache continuity. However, the current operational docs describe the cloud gateway classifying and routing each request. The public support README also describes request-level routing.
As of October 6, those descriptions disagree. I would not infer a deployed routing algorithm—or guaranteed cache behavior—from the launch headline. Link is in beta, and its docs explicitly warn that routing and model availability may change; record the version and inspect whatever serving-model and usage information your session exposes before drawing conclusions.
The current docs also put a boundary around the optional Anthropic fallback. Claude Code and Claude Desktop sessions can send difficult requests to Claude Opus when you configure an Anthropic API key, with those requests billed to that Anthropic account. Codex, OpenCode, Pi Code, and ChatGPT Desktop stay on Together models, as do Claude sessions without that optional key.
That makes “Auto” a policy choice, not just a convenience setting. Decide whether a trial may use a second provider before enabling it. For a repeatable model comparison, begin with the explicit --main route instead.
A Cost Receipt Is Not a Completed-Task Benchmark#
The usage documentation says each session prints a cost summary on exit. Claude Code’s status line also compares estimated session spend with what the same tokens would cost on Claude Opus. Cross-session usage is available with:
togetherlink usage --last 7dThat comparison is useful accounting, but it is hypothetical repricing. It does not show what Opus would actually spend solving the task: another model might use fewer turns, produce a shorter answer, or avoid a failed approach. Nor does a lower invoice measure the time a reviewer spends repairing the result.
This is the practical extension of my agent cost comparison. The denominator should be an accepted change, not a generated token. Preserving your existing login also does not move Together inference charges into a coding-agent subscription: the launch says Together usage is billed through your Together API key, and the optional Anthropic route has its own billing.
Use togetherlink models for the live lineup and rates rather than hard-coding a launch-day price table. Then run a small paired evaluation:
- Freeze the starting point. Give each run the same repository commit, task, tools, approval rules, and time limit.
- Define acceptance first. Specify the test and review requirements before either model sees the task. Include a bug fix, a dependency update, and a change that touches multiple files.
- Separate fixed-model and Auto runs. Record Link’s version, the selected route, observed serving models where available, total spend, retries, and elapsed time.
- Count failures and human repair. A cheap failed run still belongs in the total. Track review and cleanup time alongside API cost.
Keep those acceptance checks independent of the agent’s own reassurance. The warning from my AI-assisted testing analysis applies here too: a passing test generated from the same mistaken assumption is not strong evidence that the requested behavior works.
Check the Features You Actually Depend On#
“Same harness” is not a promise that every surrounding feature behaves identically. The tool-specific documentation identifies several boundaries worth testing:
| Workflow | Documented boundary | What to check in a pilot |
|---|---|---|
| OpenCode | Requires OpenCode 2; web search uses an integration connected inside OpenCode | Existing plugins, MCP tools, and search configuration |
| Pi Code | Requires 0.80.8+; current releases require Node.js 22.19+ | Extensions and the Together provider’s model selection |
| Codex CLI | Native web search remains available | Search results and the tool calls your task needs |
| ChatGPT Desktop | Native web search is disabled in the Together profile | Whether the workflow can operate without it |
Automation has another small but consequential condition: headless Claude Code runs need closed stdin, or they can wait indefinitely for more input. For example:
togetherlink --main zai-org/GLM-5.3 claude \
-p "Explain the test command for this repository; do not edit files" \
--output-format json < /dev/nullBefore scaling beyond a pilot, review the hosted providers’ data-handling terms for the code and context you will send. The docs say desktop profiles store your Together key locally; protect that profile as credential-bearing configuration. A reversible setup is useful, but it is not a privacy policy or a substitute for permission checks.
My Take#
Together Link’s strongest feature is not its savings headline. It is making the harness and the model separate purchasing decisions. If a team can trial another backend without rebuilding its tools and habits, it has a more credible alternative when price, reliability, or model behavior changes.
I would start with one terminal agent, one pinned model, and a task with clear acceptance criteria. Leave optional cross-provider routing off until you know what it contributes. The disagreement between launch and operational routing descriptions is a reason to measure the beta carefully, not a reason to assume either guaranteed savings or guaranteed failure.
My expectation is that interchangeable inference will make the quality of the surrounding workflow more important, not less. A tool that makes model switching easy should also make failed tasks and human cleanup visible. If the trial produces accepted changes with less total cost and no loss of control, expand it; if it only produces a prettier receipt, keep your current default.




