Air-Gapped AI for Defense: NexTech Solutions Deploys a VRLA Tech 4U Rackmount Workstation
NexTech Solutions LLC — a defense technology company building mission-critical communications, counter-UAS, and edge compute systems for the U.S. military and federal government — needed AI that could work with sensitive internal documents without routing anything through a commercial cloud service. Their solution: a VRLA Tech 4U Rackmount Workstation running open-weight language models on their own hardware, in their own environment, with no data leaving the building.
|
192 GB Total ECC GDDR7 VRAM |
24 cores Threadripper 9960X / 48 threads |
128 GB DDR5 ECC system memory |
The Challenge
For a defense technology company like NexTech — building edge compute systems, counter-UAS solutions, and classified communications infrastructure for military and federal clients — the risk model around AI is fundamentally different from a commercial enterprise. The same productivity tools that help a typical team draft faster and summarize longer — cloud-based AI APIs — require sending text to infrastructure the organization doesn’t own or control. For NexTech, whose work includes classified communications networks, tactical data systems, counter-drone systems, and deployable edge compute for military programs, that’s a non-starter.
The team wanted AI-assisted document drafting, summarization, and document question-and-answer — the same productivity gains available to any knowledge-work team — without the data residency exposure that comes with commercial cloud AI. That meant running capable open-weight models in-house, on hardware they control, with no telemetry and no internet dependency once deployed.
The core requirement: Run capable open-weight AI language models entirely on internal infrastructure — no cloud, no external API calls, no data leaving the controlled environment under any operating condition.
What makes this achievable today is the maturity of open-weight models. Leading models now match or exceed GPT-4-class performance on many real-world tasks — and they run entirely on hardware you own, without a network connection, with no usage data sent anywhere. The limiting factor is GPU VRAM: a 70B-parameter model at full FP16 precision requires roughly 140GB of GPU memory. That constraint drove the hardware selection.
The Build
VRLA Tech configured a 4U Rackmount Workstation built around two NVIDIA RTX PRO 6000 Blackwell Max-Q GPUs. With 96GB of ECC GDDR7 per card, the system delivers 192GB of combined GPU memory — enough to load and serve a 70B model in FP16 across both cards, or run multiple smaller models concurrently for different workflows.
This is a workstation-class platform, not a traditional server. The AMD Ryzen Threadripper 9960X on the ASRock TRX50 WS provides a workstation CPU architecture with PCIe 5.0 bandwidth and ECC memory support — the right foundation for high-VRAM GPU workloads requiring system memory integrity and sustained throughput.
| Component | Specification |
|---|---|
| System | VRLA Tech Rackmount Workstation — 4U |
| CPU | AMD Ryzen Threadripper 9960X — 24 cores / 48 threads, 4.2–5.4GHz, Zen 5 |
| CPU Cooling | 360mm AIO liquid cooler |
| Motherboard | ASRock TRX50 WS |
| System Memory | 128GB DDR5 ECC RDIMM (4×32GB, quad-channel) |
| Primary Storage | 2TB Samsung 990 Gen4 NVMe — OS and model weights |
| Secondary Storage | 8TB Seagate Exos 7200RPM SATA — data and archive |
| GPUs | 2× NVIDIA RTX PRO 6000 Blackwell Max-Q — 96GB ECC GDDR7 each (192GB total) |
| Power Supply | 1700W 80 PLUS Titanium |
| Chassis | SilverStone 4U Rackmount — rail kit included |
| Operating System | Ubuntu (pre-installed and configured) |
| Warranty & Support | 3-year parts warranty · Lifetime US-based engineer support |
192GB of combined ECC GDDR7 means the system can load and serve a 70B-parameter model at full FP16 precision — no quantization required, no performance compromise.
Storage was split across two tiers: a 2TB Gen4 NVMe for the OS and model weights (fast cold-start, models load in seconds), and an 8TB Seagate Exos for document data and archival. The Exos is engineered for 24/7 operation in always-on environments — relevant for a system that may run unattended for extended periods. The system was delivered fully configured and documented, with the inference stack set up and validated for air-gapped operation. NexTech’s team could begin working with it on day one.
What the System Runs
NexTech uses the workstation for three primary AI-assisted workflows, all running entirely within their controlled environment.
Document Drafting
Team members use the on-premise model to generate first drafts — reports, correspondence, briefing documents, and internal memos — from notes or outlines. Because the model runs locally, sensitive context about programs, clients, and contracts can be included in prompts without any external data exposure. The model retains nothing between sessions.
Summarization
Government and defense programs generate dense documentation: contracts, modification requests, technical specifications, regulatory filings, and reports. The system handles document-to-summary tasks, giving the team fast access to key points in lengthy materials. Long documents are processed entirely locally — no text or file content leaves the network.
Document Q&A
Using retrieval-augmented pipelines, the team can query against their own document library — asking plain-English questions and getting answers grounded in internal materials. This works well for navigating large contract repositories, historical correspondence, or technical specifications where the answer exists in a document but finding it manually takes too long.
Why On-Premise Made Sense Here
The conventional answer to “we need AI tools” is to subscribe to a commercial API. For most commercial teams, that’s the right call. For NexTech, it was the wrong architecture from the start.
Defense and government contracting work involves documents that contain inherently sensitive information: program details, contract values, client identities, technical specifications, and in some cases controlled unclassified information. Sending that to a third-party API — regardless of the vendor’s security posture — introduces a data residency question that has no clean answer in a regulated environment.
At the inference volume a team of knowledge workers generates, a capable on-premise workstation typically recovers its hardware cost within 12–24 months compared to commercial API spend — with no per-token cost after that and no exposure to future pricing changes.
The VRLA Tech 4U configuration gave NexTech a workstation that runs indefinitely without any network connection, generates no outbound traffic, and has no mechanism by which data could be transmitted externally — by design, not just by configuration. That’s a meaningful distinction from cloud-based deployments where network isolation is a setting that can be misconfigured.
Need an air-gapped AI workstation for a controlled environment?
VRLA Tech designs, builds, and configures 4U rackmount workstations and GPU servers for defense contractors, government programs, and regulated-industry teams that need capable AI without cloud dependency. See our defense contractor configurations. Every system ships configured, tested, and documented.
Build your air-gapped AI workstation
Configured to your environment — GPU count, CPU platform, memory, and storage. 3-year parts warranty. Lifetime US-based engineer support. Ships nationwide from Los Angeles.
Frequently Asked Questions
What AI workloads can a dual RTX PRO 6000 Blackwell rackmount workstation handle for a defense contractor?
A VRLA Tech 4U Rackmount Workstation with two NVIDIA RTX PRO 6000 Blackwell Max-Q GPUs provides 192GB of combined ECC GDDR7 VRAM — enough to run 70B parameter models in FP16 without quantization. For defense and government contractors, this supports document drafting, summarization, document Q&A, and concurrent inference across a team, all fully air-gapped with no cloud dependency. VRLA Tech builds and configures these systems in Los Angeles with a 3-year parts warranty and lifetime US-based engineer support.
Why do defense contractors run AI on-premise instead of cloud?
Defense and government contractors work with controlled, sensitive, and proprietary information that cannot be sent to a commercial cloud AI service. On-premise rackmount workstations keep all data within the organization’s controlled environment, eliminate cloud data residency risk, and ensure no telemetry or phone-home behavior. VRLA Tech configures systems specifically for air-gapped, self-contained deployment — fully documented for integration into controlled environments.
What is the difference between a rackmount workstation and a server for AI inference?
A rackmount workstation uses a workstation-class CPU platform (AMD Threadripper or Threadripper PRO) with ECC RDIMM memory, PCIe 5.0 x16 GPU slots, and workstation GPUs with active cooling — the right choice for high-VRAM AI inference and LLM serving in a rack. A traditional server platform is optimized for dense, passively cooled GPU configurations. VRLA Tech builds both — the 4U Rackmount Workstation is the recommended configuration for organizations running open-weight language models with workstation-class GPUs like the RTX PRO 6000 Blackwell.
How do I get a quote for an air-gapped AI workstation for a defense or government environment?
Contact VRLA Tech at vrlatech.com/contact-us/ or visit the defense contractor page to configure an air-gapped AI rackmount workstation for your program. VRLA Tech has built AI infrastructure for General Dynamics, Los Alamos National Laboratory, Johns Hopkins University, and government contractors including NexTech Solutions LLC. Every system ships configured and documented, with a 3-year parts warranty and lifetime US-based engineer support. Building custom workstations and GPU systems in Los Angeles since 2016.
Built by the VRLA Tech engineering team in Los Angeles. VRLA Tech has been building custom AI workstations and GPU servers for defense contractors, researchers, and enterprise teams since 2016.




