- Ubuntu Server 24.04 LTS
- 4 vCPU, 8 GiB of memory, 64 GiB of system storage
- Own identity from a full clone of the Ubuntu Server template
- Cloud-Init, SSH and guest agent validated
- OpenCode 1.18.25
- Hermes Agent 0.20.6
- No local model or GPU
- OpenCode installation and version checked
- Hermes installation checked via its built-in diagnostics
- Ollama's OpenAI-compatible endpoint reached from VM301
- Model
qwen3.6:35bdetected by the client - Direct generation succeeded
- Real OpenCode generation succeeded with the same model
- Real Hermes generation succeeded with
qwen3.6:35b— served context of 65,536 tokens, deliberately restricted tool set - OpenCode benchmark succeeded on Python, HTML and n8n JSON tasks (review and targeted fixes where needed)
This validation demonstrates the full chain between an isolated development agent and GPU-accelerated local inference.
VM301
├── Hermes Agent
├── OpenCode
└── API d'inférence autorisée sur le réseau privé
└── Ollama sur la VM GPU
└── modèles locaux chargés sur la RTX 5090The client VM stays lightweight. Model storage, loading into video memory and inference optimizations remain centralized on the GPU node. The API is not published directly on the Internet and access is limited to explicitly authorized clients.
qwen3.6:35b is the validated reference for Hermes, with a context of 65,536 tokens.
qwen3-coder:30b is validated with OpenCode for development tasks: Python succeeded directly, while HTML and n8n JSON each required a targeted fix before final validation.
- Generation and maintenance of Python code
- Creation of HTML pages and components
- Production of structured JSON data
- Preparation of n8n workflows
- Repository analysis and multi-file changes
- Writing technical documentation
Generated n8n workflows remain subject to human validation: syntax check, import disabled and review before any activation, especially in the presence of expressions, network calls or credentials.
- No inference port directly published on the Internet
- Explicit filtering of client machines
- Client configuration without a token when the private endpoint does not require one
- Secrets, addresses and internal rules excluded from the public source
- Future remote access planned via an authenticated private tunnel, not by opening the Ollama port
- Reproducible VM creation from a sanitized template
- Installation and diagnosis of recent agentic clients
- Integration of a self-hosted OpenAI-compatible API
- Separation between the agentic workstation and GPU compute
- Network filtering following the least-privilege principle
- Progressive validation of a full LLM chain
- Documented distinction between validated state and ongoing experimentation
- VM301's application persistence still needs to be checked after a reboot;
- the Windows development station is documented separately on the VM302 page, with its own validations from September 20, 2026 — these do not validate VM301's reboot recovery;
- remote access for client workstations will need to go through a dedicated, validated VPN architecture.