Air-Gapped AI for Defense: NexTech Solutions Deploys a VRLA Tech 4U Rackmount Workstation

Air-Gapped AI for Defense: NexTech Solutions Deploys a VRLA Tech 4U Rackmount Workstation

NexTech Solutions LLC — a defense technology company building mission-critical communications, counter-UAS, and edge compute systems for the U.S. military and federal government — needed AI that could work with sensitive internal documents without routing anything through a commercial cloud service. Their solution: a VRLA Tech 4U Rackmount Workstation running open-weight language models on their own hardware, in their own environment, with no data leaving the building.

192 GB
Total ECC GDDR7 VRAM
24 cores
Threadripper 9960X / 48 threads
128 GB
DDR5 ECC system memory

The Challenge

For a defense technology company like NexTech — building edge compute systems, counter-UAS solutions, and classified communications infrastructure for military and federal clients — the risk model around AI is fundamentally different from a commercial enterprise. The same productivity tools that help a typical team draft faster and summarize longer — cloud-based AI APIs — require sending text to infrastructure the organization doesn’t own or control. For NexTech, whose work includes classified communications networks, tactical data systems, counter-drone systems, and deployable edge compute for military programs, that’s a non-starter.

The team wanted AI-assisted document drafting, summarization, and document question-and-answer — the same productivity gains available to any knowledge-work team — without the data residency exposure that comes with commercial cloud AI. That meant running capable open-weight models in-house, on hardware they control, with no telemetry and no internet dependency once deployed.

The core requirement: Run capable open-weight AI language models entirely on internal infrastructure — no cloud, no external API calls, no data leaving the controlled environment under any operating condition.

What makes this achievable today is the maturity of open-weight models. Leading models now match or exceed GPT-4-class performance on many real-world tasks — and they run entirely on hardware you own, without a network connection, with no usage data sent anywhere. The limiting factor is GPU VRAM: a 70B-parameter model at full FP16 precision requires roughly 140GB of GPU memory. That constraint drove the hardware selection.


The Build

VRLA Tech configured a 4U Rackmount Workstation built around two NVIDIA RTX PRO 6000 Blackwell Max-Q GPUs. With 96GB of ECC GDDR7 per card, the system delivers 192GB of combined GPU memory — enough to load and serve a 70B model in FP16 across both cards, or run multiple smaller models concurrently for different workflows.

This is a workstation-class platform, not a traditional server. The AMD Ryzen Threadripper 9960X on the ASRock TRX50 WS provides a workstation CPU architecture with PCIe 5.0 bandwidth and ECC memory support — the right foundation for high-VRAM GPU workloads requiring system memory integrity and sustained throughput.

192 GB
GPU VRAM total
2× 96 GB
ECC GDDR7 per GPU
70B+
Model params at FP16
Air-gap
No cloud dependency
ComponentSpecification
SystemVRLA Tech Rackmount Workstation — 4U
CPUAMD Ryzen Threadripper 9960X — 24 cores / 48 threads, 4.2–5.4GHz, Zen 5
CPU Cooling360mm AIO liquid cooler
MotherboardASRock TRX50 WS
System Memory128GB DDR5 ECC RDIMM (4×32GB, quad-channel)
Primary Storage2TB Samsung 990 Gen4 NVMe — OS and model weights
Secondary Storage8TB Seagate Exos 7200RPM SATA — data and archive
GPUs2× NVIDIA RTX PRO 6000 Blackwell Max-Q — 96GB ECC GDDR7 each (192GB total)
Power Supply1700W 80 PLUS Titanium
ChassisSilverStone 4U Rackmount — rail kit included
Operating SystemUbuntu (pre-installed and configured)
Warranty & Support3-year parts warranty · Lifetime US-based engineer support

192GB of combined ECC GDDR7 means the system can load and serve a 70B-parameter model at full FP16 precision — no quantization required, no performance compromise.

Storage was split across two tiers: a 2TB Gen4 NVMe for the OS and model weights (fast cold-start, models load in seconds), and an 8TB Seagate Exos for document data and archival. The Exos is engineered for 24/7 operation in always-on environments — relevant for a system that may run unattended for extended periods. The system was delivered fully configured and documented, with the inference stack set up and validated for air-gapped operation. NexTech’s team could begin working with it on day one.


What the System Runs

NexTech uses the workstation for three primary AI-assisted workflows, all running entirely within their controlled environment.

Document Drafting

Team members use the on-premise model to generate first drafts — reports, correspondence, briefing documents, and internal memos — from notes or outlines. Because the model runs locally, sensitive context about programs, clients, and contracts can be included in prompts without any external data exposure. The model retains nothing between sessions.

Summarization

Government and defense programs generate dense documentation: contracts, modification requests, technical specifications, regulatory filings, and reports. The system handles document-to-summary tasks, giving the team fast access to key points in lengthy materials. Long documents are processed entirely locally — no text or file content leaves the network.

Document Q&A

Using retrieval-augmented pipelines, the team can query against their own document library — asking plain-English questions and getting answers grounded in internal materials. This works well for navigating large contract repositories, historical correspondence, or technical specifications where the answer exists in a document but finding it manually takes too long.


Why On-Premise Made Sense Here

The conventional answer to “we need AI tools” is to subscribe to a commercial API. For most commercial teams, that’s the right call. For NexTech, it was the wrong architecture from the start.

Defense and government contracting work involves documents that contain inherently sensitive information: program details, contract values, client identities, technical specifications, and in some cases controlled unclassified information. Sending that to a third-party API — regardless of the vendor’s security posture — introduces a data residency question that has no clean answer in a regulated environment.

At the inference volume a team of knowledge workers generates, a capable on-premise workstation typically recovers its hardware cost within 12–24 months compared to commercial API spend — with no per-token cost after that and no exposure to future pricing changes.

The VRLA Tech 4U configuration gave NexTech a workstation that runs indefinitely without any network connection, generates no outbound traffic, and has no mechanism by which data could be transmitted externally — by design, not just by configuration. That’s a meaningful distinction from cloud-based deployments where network isolation is a setting that can be misconfigured.

Need an air-gapped AI workstation for a controlled environment?

VRLA Tech designs, builds, and configures 4U rackmount workstations and GPU servers for defense contractors, government programs, and regulated-industry teams that need capable AI without cloud dependency. See our defense contractor configurations. Every system ships configured, tested, and documented.

Talk to an engineer →


Build your air-gapped AI workstation

Configured to your environment — GPU count, CPU platform, memory, and storage. 3-year parts warranty. Lifetime US-based engineer support. Ships nationwide from Los Angeles.

See defense configurations →

Ready to buy?

Frequently Asked Questions

What AI workloads can a dual RTX PRO 6000 Blackwell rackmount workstation handle for a defense contractor?

A VRLA Tech 4U Rackmount Workstation with two NVIDIA RTX PRO 6000 Blackwell Max-Q GPUs provides 192GB of combined ECC GDDR7 VRAM — enough to run 70B parameter models in FP16 without quantization. For defense and government contractors, this supports document drafting, summarization, document Q&A, and concurrent inference across a team, all fully air-gapped with no cloud dependency. VRLA Tech builds and configures these systems in Los Angeles with a 3-year parts warranty and lifetime US-based engineer support.

Why do defense contractors run AI on-premise instead of cloud?

Defense and government contractors work with controlled, sensitive, and proprietary information that cannot be sent to a commercial cloud AI service. On-premise rackmount workstations keep all data within the organization’s controlled environment, eliminate cloud data residency risk, and ensure no telemetry or phone-home behavior. VRLA Tech configures systems specifically for air-gapped, self-contained deployment — fully documented for integration into controlled environments.

What is the difference between a rackmount workstation and a server for AI inference?

A rackmount workstation uses a workstation-class CPU platform (AMD Threadripper or Threadripper PRO) with ECC RDIMM memory, PCIe 5.0 x16 GPU slots, and workstation GPUs with active cooling — the right choice for high-VRAM AI inference and LLM serving in a rack. A traditional server platform is optimized for dense, passively cooled GPU configurations. VRLA Tech builds both — the 4U Rackmount Workstation is the recommended configuration for organizations running open-weight language models with workstation-class GPUs like the RTX PRO 6000 Blackwell.

How do I get a quote for an air-gapped AI workstation for a defense or government environment?

Contact VRLA Tech at vrlatech.com/contact-us/ or visit the defense contractor page to configure an air-gapped AI rackmount workstation for your program. VRLA Tech has built AI infrastructure for General Dynamics, Los Alamos National Laboratory, Johns Hopkins University, and government contractors including NexTech Solutions LLC. Every system ships configured and documented, with a 3-year parts warranty and lifetime US-based engineer support. Building custom workstations and GPU systems in Los Angeles since 2016.


Built by the VRLA Tech engineering team in Los Angeles. VRLA Tech has been building custom AI workstations and GPU servers for defense contractors, researchers, and enterprise teams since 2016.

Leave a Reply

Your email address will not be published. Required fields are marked *

NOTIFY ME We will inform you when the product arrives in stock. Please leave your valid email address below.
U.S Based Support
Based in Los Angeles, our U.S.-based engineering team supports customers across the United States, Canada, and globally. You get direct access to real engineers, fast response times, and rapid deployment with reliable parts availability and professional service for mission-critical systems.
Expert Guidance You Can Trust
Companies rely on our engineering team for optimal hardware configuration, CUDA and model compatibility, thermal and airflow planning, and AI workload sizing to avoid bottlenecks. The result is a precisely built system that maximizes performance, prevents misconfigurations, and eliminates unnecessary hardware overspend.
Reliable 24/7 Performance
Every system is fully tested, thermally validated, and burn-in certified to ensure reliable 24/7 operation. Built for long AI training cycles and production workloads, these enterprise-grade workstations minimize downtime, reduce failure risk, and deliver consistent performance for mission-critical teams.
Future Proof Hardware
Built for AI training, machine learning, and data-intensive workloads, our high-performance workstations eliminate bottlenecks, reduce training time, and accelerate deployment. Designed for enterprise teams, these scalable systems deliver faster iteration, reliable performance, and future-ready infrastructure for demanding production environments.
Engineers Need Faster Iteration
Slow training slows product velocity. Our high-performance systems eliminate queues and throttling, enabling instant experimentation. Faster iteration and shorter shipping cycles keep engineers unblocked, operating at startup speed while meeting enterprise demands for reliability, scalability, and long-term growth today globally.
Cloud Cost are Insane
Cloud GPUs are convenient, until they become your largest monthly expense. Our workstations and servers often pay for themselves in 4–8 weeks, giving you predictable, fixed-cost compute with no surprise billing and no resource throttling.