Standard coding assistants lack native platform awareness. They cannot trigger Domino jobs, query workspace environments, register models, or enforce governance rules. Domino Skills bridges this context gap for tools like Claude Code, but teams using custom harnesses need that same awareness decoupled from model choice.
The solution
Cline is an Apache 2.0 licensed coding agent running as a VS Code extension. It treats the underlying model as a configuration setting, enabling seamless switching between Anthropic, AWS Bedrock, Google Vertex AI, OpenAI-compatible endpoints, or local runtimes like Ollama and LM Studio.
Cline open model selection
Integrating a Cline skill pack with an MCP server wrapping the Domino REST API ensures platform awareness remains independent of model selection. This combination allows developers to run a coding agent against a Domino-hosted open-source model while keeping inference on internal infrastructure, route the same agent through the Domino AI Gateway for centralized authentication and logging, submit and monitor Domino Jobs directly from VS Code, and swap models mid-project without losing platform context.
When to consider Cline in Domino
External Inference Constraints: When security mandates specific providers like AWS Bedrock or Google Vertex AI, Cline accesses these through a unified settings panel, preventing vendor lock-in from hindering agentic development.
Infrastructure Isolation: Air-gapped deployments and strict data residency requirements demand internal inference. A Domino-hosted model behind an OpenAI-compatible endpoint satisfies these constraints, allowing Cline to connect as a first-class provider.
Model Evaluation: Running identical scaffolding tasks across different models yields better insights than synthetic benchmarks. Cline allows developers to seamlessly switch between frontier models, hosted open-source models, and local runtimes to evaluate performance on real workloads.
Cost Optimization: High-volume tasks requiring type annotations, docstring generation, or test scaffolding can be routed to cost-effective models, reserving frontier models for complex architectural design.
Security Reviews: Security audits are simplified by Cline's Apache 2.0 license and open-source nature. Security teams can audit the agent's file and credential access patterns directly.
Existing Open Source LLMs: Models deployed in Domino for application serving can concurrently back a coding agent on the same endpoint without additional infrastructure overhead.
How to set up Cline with Domino context
You will need VS Code with the Cline extension, uv to run the MCP server, Python 3.11 or later, and a Domino account with API access. Run the steps below from the repository root, since the symlink commands resolve relative to your current directory.
Install the skills
Cline reads skills from ~/.cline/skills/ globally and .cline/skills/ per project. Each skill is a SKILL.md file with name and description frontmatter, and Cline routes to them automatically by matching the description against what you asked for. You can also call one directly with /skillname when you want to force it.
Porting Domino Skills to Cline means editing the frontmatter, the YAML metadata block at the top of the file. Both formats are markdown with name and description at the top. The Claude Code subagent fields (tools:, model:, skills:) have no Cline equivalent and get dropped, and the skill body carries over unchanged.
Install the workflows
Workflows are the explicit counterpart to skills. Where a skill loads when its description matches, a workflow is a multi-step procedure you invoke by name, and nothing triggers it on its own. Install globally to ~/Documents/Cline/Workflows/, or per project to .clinerules/workflows/.
The rule file is standing context that loads in every conversation. It tells Cline to resolve which Domino instance it is talking to before assuming one, and to check the live swagger and the platform docs (docs.domino.ai for Domino Cloud, the matching version at docs.dominodatalab.com for self-managed) before implementing against an endpoint. It also prevents the failure mode where a model gets stuck in a long debug loop built on an endpoint the agent invented.
Cline supports both local stdio servers and remote Streamable HTTP servers. The Domino REST API server runs locally under uv, which provisions its own environment from the lockfile on first use.
Credentials live in ~/.domino/.env, next to your existing Domino CLI data. A single instance needs DOMINO_HOST and DOMINO_API_KEY. Several instances use an alias list, which is the common case once you have a dev cluster and a production one.
Bash / Shell
DOMINO_CLUSTERS=prod,dev
DOMINO_HOST_PROD=https://your-domino-host
DOMINO_API_KEY_PROD=<key>
DOMINO_HOST_DEV=https://your-dev-domino-host
DOMINO_API_KEY_DEV=<key>
Every tool on the server takes an optional cluster argument matching one of those aliases, and a list_domino_clusters tool reports what is configured.
Watch the credential path here, because it is where a ported skill breaks first. DOMINO_API_HOST and the workspace bearer token exist inside a Domino workspace, job, or app. Cline runs on your laptop, where neither is set, so a skill that was written for in-workspace use will fail when the same code runs from VS Code. Resolve the host through list_domino_clusters and authenticate with the X-Domino-Api-Key header instead.
What to consider when choosing a model
Cline exposes model choice as a provider selection plus a model name. The options below cover most enterprise cases.
Option
How Cline reaches it
Use it when
Frontier model, vendor API
Anthropic, OpenAI, or Google provider with an API key
Quality is the constraint and sending code to a vendor API is allowed
Frontier model, cloud marketplace
AWS Bedrock or Google Vertex AI provider
Your contract, billing, and data-residency terms already run through that cloud
Domino-hosted open source model
OpenAI Compatible provider, base URL set to your model endpoint
Inference has to stay on infrastructure you control
Local model
Ollama or LM Studio provider
You are offline, air-gapped, or iterating on cheap high-volume work
Agentic coding demands more from a model than chat does. The agent has to call tools correctly, read the results, and recover when a command fails, and small models degrade on that long before they degrade on writing a function. If you are evaluating a hosted open-source model for this, test it on a multi-step task with real tool calls rather than a code-completion prompt.
Model choice is not permanent. Nothing in the skills, workflows, rules, or MCP server is tied to a provider. Swap the model and you change what does the reasoning, not what the agent knows about Domino.
Pointing Cline at a Domino-hosted model
Domino lets you deploy open-source LLMs as production endpoints on your own infrastructure, served through an optimized vLLM runtime behind an OpenAI-compatible API. Cline's OpenAI Compatible provider takes a base URL, an API key, and a model ID, and it does not check who is on the other end as long as the responses match the schema.
In Cline's settings, select OpenAI Compatible as the provider and set the base URL to your model endpoint:
https://your-domino-host/model-endpoint/v1
Set the model ID to the name your endpoint serves and the API key to your Domino credential. Cline handles the agent loop, the file edits, the tool calls, and the context management, while the weights run on your hardware.
Size the endpoint for the workload before you point an agent at it. A 32-billion-parameter model in 4-bit quantization needs roughly 32 billion x 0.5 bytes = 16 GB for weights alone, before KV cache and runtime overhead, and a coding agent runs far longer contexts than a chat client does, which makes the KV cache the line item that surprises people. The Deploying Self-Hosted LLMs Blueprint works through the sizing math and the vLLM configuration in detail.
Routing through Domino AI Gateway
Pointing at a model endpoint solves where inference happens. It does not give you a record of what was sent. What gives you that depends on your deployment type. The LLM gateway is the proxy layer that does: it centralizes authentication so individual API keys stop circulating, enforces access controls on which models a project can reach, and logs every LLM interaction for audit.
Set the OpenAI Compatible base URL to your gateway endpoint rather than the model endpoint, and Cline's traffic joins everything else running through it. Behind the gateway you can route to a vendor API, a cloud marketplace model, or a model you host, and change that routing without touching a single developer's settings.
If the agent is writing code that will handle regulated data, the gateway is the piece that makes the coding step auditable alongside the rest of the workflow.
Running a model on your laptop
Cline supports Ollama and LM Studio directly, which is the fastest way to get an agent running with no network egress at all. It suits an air-gapped workstation, or a repetitive refactor you would rather not meter against an API.
However, know the tradeoffs you are making. A model that fits in consumer VRAM will be slower per token and less reliable on long tool-calling chains than a hosted 70-billion-parameter model or a frontier API. Pick a coder-tuned model with documented tool-calling support, give it small scoped tasks, and keep the plan-then-act discipline. Cline's plan mode helps here: it restricts the agent to reading and analysis with no file writes or command execution, so you can check the approach before a weaker model starts editing.
Administering the underlying Domino deployment
Everything above treats Domino as a platform you build on. The same plugin also carries the skills a platform team uses to run the deployment itself: checking whether a given credential set can actually reach a cluster, getting a kubectl session through Teleport, and inspecting the AWS resources (EKS, S3, EFS, IAM, CloudWatch) that Domino's control plane and compute actually run on. None of this requires the coding-assistant framing above; it's the same install, pointed at operations instead of application code.
The three-leg access check. The domino-access skill verifies cluster access in three separate legs, run in a fixed order, rather than stopping at the first success:
REST API: check_domino_api_access against the cluster alias, via the same MCP server used for jobs and environments. A 401/403 means the API key needs regenerating; a timeout means the cluster itself is unreachable from here.
Teleport and kubectl: scripts/domino-tsh login <alias>, then kubectl cluster-info to confirm the context actually resolves.
AWS: aws sts get-caller-identity confirms the session is still active under Domino's Okta-federated admin role.
Each leg fails independently and for a different reason, so the skill reports all three rather than guessing which symptom points to which one.
Reaching the Kubernetes layer. Domino's EKS clusters aren't reached with kubectl or aws eks update-kubeconfig directly. Access goes through Teleport, and Domino runs more than one Teleport major version at once (an older dev instance, a current production one). The domino-teleport skill's scripts/domino-tsh picks a compatible tsh binary for whichever cluster you're targeting automatically, rather than assuming one is installed under a particular name. See the Teleport version dispatch flow figure at the end of this section for how that selection works.
AWS-level operations. Once you're authenticated (okta-aws, an interactive Okta OIDC flow into Domino's admin role), the aws-ops skill covers the CLI patterns for inspecting the infrastructure underneath a cluster:
EKS: aws eks describe-cluster and list-nodegroups for control-plane and node-group state (the workload-facing kubectl access above is separate from this).
S3 and EFS: bucket and file-system inspection for the storage backing Domino Datasets and project volumes.
IAM: auditing what a cluster's execution roles can actually do, for example aws iam list-attached-role-policies --role-name <role> when tracking down why an IRSA-bound service account can or can't reach a bucket, or aws iam get-role when confirming a trust policy matches what Domino's operator expects.
CloudWatch Logs: aws logs tail <log-group> --follow for control-plane and platform-service logs that don't show up in a project's own job/app logs.
aws-ops is deliberately incomplete. Its own instructions tell the agent to append a new section documenting any AWS service it handles that isn't covered yet, so the skill accumulates real command patterns and Domino-specific conventions (bucket naming, IRSA role names, that kind of thing) as your team actually uses it, instead of staying frozen at whatever shipped.
Teleport Version Dispatch Flow
Extending the pack
The extension points are independent of each other, so you can add any one of them without touching the rest.
Skills carry most of the weight. A skill is a markdown file with a name and a description, and Cline loads its full contents when the description matches the task. Good candidates are the conventions you would have to tell a new team member: your hardware tier naming, your MLflow experiment tagging standard, your data source connection patterns, your app deployment checklist. Keep each one focused and move long reference material into sibling files, since only the description is loaded up front.
A skill can also instruct the agent to append to its own file whenever it handles a case the file does not yet cover. An operations skill written that way documents each new service the first time someone uses it, so the file grows with the team instead of freezing at whatever shipped.
Workflows are for procedures with a fixed sequence. Project scaffolding, a debug protocol, a tracing setup that always runs the same five steps. They only run when you call them by name, so they suit work you want deliberate rather than automatic.
Rules are standing instructions that load every conversation. Use them for discipline that should never depend on the model noticing it applies, like verifying against live docs before implementing, or confirming which cluster is the target before running anything.
Hooks are the deterministic layer. Where a rule is an instruction the model can miss, a hook is a script that runs whether or not the model cooperates. Cline runs one executable per event, named for the event and placed in ~/Documents/Cline/Hooks/ globally or .clinerules/hooks/ per project, with no per-tool matcher configuration. Cline's hooks page is not published documentation, so this list comes from reading the extension's source directly, not from Cline's docs. Four events cover the Domino use cases below: PreToolUse, PostToolUse, SessionStart, and UserPromptSubmit. Three more exist (PreCompact, PostToolUseFailure, PostToolBatch) without an obvious Domino use case yet.
Event
Fires when
Domino use case
PreToolUse
Before a tool runs
Block writes to protected project paths
PostToolUse
After a tool completes
Check that an app binds to 0.0.0.0, run a formatter, lint a Dockerfile
SessionStart
A session begins
Load Domino credentials and inject current project context
UserPromptSubmit
A prompt is submitted
Prepend the target cluster and environment to every request
Because there is one file per event, everything you want checked after a write goes in the same script and dispatches internally on the tool name and file path. Have it exit 0 unless the check is one you want to block on.
What does not port
Cline has no sub-agent mechanism, so anything built as an isolated agent context in the Claude Code plugin becomes an ordinary auto-routed skill. The knowledge carries over intact. What you lose is the separate context window, which means a long debugging session shares context with everything else in the conversation.
Cline also has no swappable output style or persona mechanism. Claude Code's output styles append a formatted block after each task, an explainer for onboarding or a checklist for deployment review. The closest equivalent is a global rule you toggle in Cline's Rules panel, which reproduces the content without the one-click switch. Cline also has no Stop-equivalent hook event, so the Ralph Loop pattern from the Claude Code plugin, running a task autonomously until a completion condition is met, has no direct port. Cline's verified hook events cover before and after a tool call, session start, and prompt submission, not session end.
Where to go next
If your model choice is Claude and you want the agent running inside a Domino workspace rather than on your laptop, Claude Code on Domino covers the pre-installed path, the Domino Skills framework, and the environment configuration for self-managed deployments.
For the model side, Deploying Self-Hosted LLMs in Domino covers registering an open source model, sizing the hardware tier, and configuring vLLM for the endpoint this blueprint points Cline at.
Michael Snyder is a senior data architect and AI/ML systems leader with over a decade of experience building enterprise-scale data platforms and analytics for mission-critical defense and logistics environments. He currently serves as a Forward Deployed Engineer at Domino Data Lab, supporting public sector customers on AI/ML platform deployment and applied machine learning initiatives, including automated target recognition systems for Navy drone programs.