Deployment Planning
Prepare hardware, networking, models, and operational ownership
Table of Contents
About this guide
Agree on the deployment configuration with Ask Sage before installation. Model size, simultaneous users, document workloads, network restrictions, and recovery requirements determine the resources you need.
Choose the deployment type
| Deployment | What to plan for |
|---|---|
| Appliance | The supplied hardware, GPU configuration, storage, power, cooling, and network connections. Follow its desktop, rack, or rugged installation requirements. |
| Multi-node appliance | The required nodes, interconnects, and placement of application and model services. Each node has its own resource budget. |
| Customer-hosted Kubernetes | Your infrastructure team provides and operates the cluster, GPU resources, storage, ingress, and infrastructure recovery. Use the Manager for application and model operations. |
Do not size a model by adding together the memory printed on every GPU. A model must fit the resources assigned to it, including runtime memory. Spreading a model across GPUs requires a supported model, serving engine, and interconnect configuration. See model management.
Prepare the site
| Area | Information to confirm |
|---|---|
| Physical installation | Power, cooling, placement or rack space, supplied cables, and physical access for maintenance. |
| Network | Appliance addresses or DHCP reservations, subnet, gateway, DNS, and the user networks allowed to connect. |
| Names and certificates | Workspace, API, and Manager addresses; certificate issuer, trust chain, and renewal owner. |
| Connectivity | Fully air-gapped, periodically connected, or connected; approved destinations and proxy requirements. |
| Identity | Initial administrator, user provisioning, local login or SSO, and a recovery access method. |
| Models | Approved models, licenses, model files, serving-engine compatibility, and expected concurrency. |
| Data | Storage capacity, retention, dataset access rules, backup scope, and restore requirements. |
| Maintenance | Release source, media transfer process, maintenance windows, security review, and support contact. |
For customer-hosted infrastructure, also confirm the Kubernetes and GPU-driver versions, persistent storage, supported CPU architecture, image registry, and installation privileges with the deployment team. Use installation media matched to the hardware architecture and deployment configuration.
Installation and handover
A delivered appliance can arrive with Ask Sage installed. Follow Appliance kit setup for unpacking and cabling a supplied portable kit, then Getting started for sign-in and acceptance checks. A bare server or customer-hosted cluster requires the installation package and runbook supplied for that deployment; connecting power alone does not install Ask Sage.
For a disconnected installation, the deployment team must stage the software, container images, model files, and other dependencies before entering the isolated network. Use the supplied release and media instructions rather than downloading replacement components individually.
Before handover is complete, record:
- Platform, installed release, node inventory, and application addresses.
- Named owners for appliance administration, identity, networking, updates, and recovery.
- Successful sign-in, local chat, and document upload/retrieval tests.
- Available models and the expected behavior after restart.
- Location of protected recovery material and the agreed backup and restore procedure.
- Known limitations, security findings, and any services that still require connectivity.
Store passwords and keys in your organization’s credential system, separately from the general handover record.