VM301 is a Linux virtual machine dedicated to development assistants. It runs Hermes Agent and OpenCode while delegating inference to the platform's GPU VM. This separation provides an isolated agentic environment without duplicating models or using an additional GPU.
✅ Deployed base
  • Ubuntu Server 24.04 LTS
  • 4 vCPU, 8 GiB of memory, 64 GiB of system storage
  • Own identity from a full clone of the Ubuntu Server template
  • Cloud-Init, SSH and guest agent validated
  • OpenCode 1.18.25
  • Hermes Agent 0.20.6
  • No local model or GPU
🔎 Validation performed
  • OpenCode installation and version checked
  • Hermes installation checked via its built-in diagnostics
  • Ollama's OpenAI-compatible endpoint reached from VM301
  • Model qwen3.6:35b detected by the client
  • Direct generation succeeded
  • Real OpenCode generation succeeded with the same model
  • Real Hermes generation succeeded with qwen3.6:35b — served context of 65,536 tokens, deliberately restricted tool set
  • OpenCode benchmark succeeded on Python, HTML and n8n JSON tasks (review and targeted fixes where needed)

This validation demonstrates the full chain between an isolated development agent and GPU-accelerated local inference.

🧭 Architecture — isolated agent, centralized compute
separation between the agentic workstation and the GPU node
VM301
├── Hermes Agent
├── OpenCode
└── API d'inférence autorisée sur le réseau privé
    └── Ollama sur la VM GPU
        └── modèles locaux chargés sur la RTX 5090

The client VM stays lightweight. Model storage, loading into video memory and inference optimizations remain centralized on the GPU node. The API is not published directly on the Internet and access is limited to explicitly authorized clients.

🎛️ Model choice

qwen3.6:35b is the validated reference for Hermes, with a context of 65,536 tokens.

qwen3-coder:30b is validated with OpenCode for development tasks: Python succeeded directly, while HTML and n8n JSON each required a targeted fix before final validation.

🧰 Intended uses
  • Generation and maintenance of Python code
  • Creation of HTML pages and components
  • Production of structured JSON data
  • Preparation of n8n workflows
  • Repository analysis and multi-file changes
  • Writing technical documentation

Generated n8n workflows remain subject to human validation: syntax check, import disabled and review before any activation, especially in the presence of expressions, network calls or credentials.

🔐 Security
  • No inference port directly published on the Internet
  • Explicit filtering of client machines
  • Client configuration without a token when the private endpoint does not require one
  • Secrets, addresses and internal rules excluded from the public source
  • Future remote access planned via an authenticated private tunnel, not by opening the Ollama port
🎯 Skills demonstrated
  • Reproducible VM creation from a sanitized template
  • Installation and diagnosis of recent agentic clients
  • Integration of a self-hosted OpenAI-compatible API
  • Separation between the agentic workstation and GPU compute
  • Network filtering following the least-privilege principle
  • Progressive validation of a full LLM chain
  • Documented distinction between validated state and ongoing experimentation
🚧 Limits & open work:

Related pages

References & Sources

CategoryResourceURL
Official documentationOllama — Modèles & API compatible OpenAIgithub.com/ollama/ollama
Official documentationUbuntu Server 24.04 LTSubuntu.com/server
Agentic clientsOpenCode 1.18.25 · Hermes Agent 0.20.6 — versions relevées sur la VM—
ValidationCréation et identité de VM301 contrôlées · endpoint Ollama atteint · génération Hermes avec qwen3.6:35b · benchmark OpenCode avec qwen3-coder:30b (Python/HTML/JSON n8n), 2026-08-31—
Content of this pageShared under CC BY-SA 4.0creativecommons.org/licenses/by-sa/4.0