AI Hardware Compatibility Guide · MAQ Buying Guide

Which AI model do you want to run,
and how should you configure for it?

From gemma4-26b, gpt-oss-120b, Qwen3, to AI Agent workflows — we lay out "model → VRAM → system memory → CPU → recommended build" for every case, each mapped to an actual MAQ machine with an online quote, and VRAM figures always state the quantization used.

Get an online quote → Ask on LINE
MAQ AI workstation: a hardware environment for running Ollama and Docker locally
32GB
single GPU, 4-bit 26B
96GB
single GPU, gpt-oss-120B
4-way
multi-GPU scaling
34Kfrom
AI Agent PC

How to read this table

Start with your workload (inference / fine-tuning / image generation / AI Agent), then match model size to the required VRAM and system memory. The right column shows the MAQ machine already configured for that need, with a link to an instant online quote. VRAM shifts with quantization (4-bit / fp8 / fp16) and context length — figures below are conservative estimates.

Entry to Mid-Size LLM Local Inference (7B–32B)

This tier is the "one card is enough" sweet spot: at 4-bit quantization, 7B–14B models need roughly 5–12GB VRAM and 32B needs about 18–24GB — a single 24–32GB professional card runs them smoothly. What matters is not raw card size but whether VRAM covers "quantized weights + KV cache" (long context can eat several extra GB); moving to full precision (FP16) or extending context can double to quadruple the requirement. (Models shown reflect mainstream 2026 releases; newer versions at the same tier have similar hardware needs.)

Use case · representative models Recommended spec Key differences Recommended MAQ build
Entry-level desktop assistant / RAG Q&A
Llama 3.1 8B · Gemma 3 12B · Qwen3 8B
5–12GB (4-bit); FP16: 8B/9B ≈16–18GB, 14B ≈28GB
32GB DDR5
12-core Ryzen 9 (9900X3D class)
At 4-bit, 8B–14B models run comfortably, with 24GB still leaving room for 8K–32K context KV cache; for full-precision 14B or running multiple models concurrently, go straight to a 32GB machine.
AI-Eco (RTX PRO 4000 24GB)
約 10 萬
Get quote →
Mid-size production inference: gpt-oss-20b online service
gpt-oss-20b · Qwen3 14B · Gemma 3 12B
16–24GB (gpt-oss-20b ≈16GB quantized; 14B Q8 ≈15GB)
64GB DDR5
12-core Ryzen 9 9900X3D
Pairs Ryzen 9 9900X3D with an AMD Radeon AI PRO R9700 32GB GPU. 32GB VRAM fully loads gpt-oss-20b (AMD platforms use ROCm's FP8/quantized path; native MXFP4 acceleration is exclusive to NVIDIA Blackwell machines), with headroom left for concurrency and long context — ready out of the box.
AI-Medium (AMD Radeon AI PRO R9700 32GB GPU, preloaded with gpt-oss-20b)
約 16 萬
Get quote →
32B-class high-quality inference (coding assistance / advanced Q&A)
Qwen3 32B · Gemma 3 27B
18–24GB (4-bit); Q8: Gemma 27B ≈27–30GB (fits a single 32GB card), Qwen 32B ≈34–35GB (slightly over 32GB)
64GB DDR5
16-core Ryzen 9 9950X3D
At 4-bit, a single 32GB card runs it with room for context; Gemma 27B fits even at Q8, while Qwen 32B at Q8 slightly exceeds 32GB and needs a shorter sequence or 48GB. The difference from AI-Medium (NT$151K / Radeon AI PRO R9700 · ROCm): NVIDIA's CUDA ecosystem, long-run-stable Pro drivers, and Gemma preloaded.
AI-Medium-Gemma (RTX PRO 4500 32GB)
約 27 萬
Get quote →
Entry-level desktop assistant / RAG Q&A
約 10 萬
Llama 3.1 8B · Gemma 3 12B · Qwen3 8B
5–12GB (4-bit); FP16: 8B/9B ≈16–18GB, 14B ≈28GB
32GB DDR5・12-core Ryzen 9 (9900X3D class)
At 4-bit, 8B–14B models run comfortably, with 24GB still leaving room for 8K–32K context KV cache; for full-precision 14B or running multiple models concurrently, go straight to a 32GB machine.
AI-Eco (RTX PRO 4000 24GB)
Get quote →
Mid-size production inference: gpt-oss-20b online service
約 16 萬
gpt-oss-20b · Qwen3 14B · Gemma 3 12B
16–24GB (gpt-oss-20b ≈16GB quantized; 14B Q8 ≈15GB)
64GB DDR5・12-core Ryzen 9 9900X3D
Pairs Ryzen 9 9900X3D with an AMD Radeon AI PRO R9700 32GB GPU. 32GB VRAM fully loads gpt-oss-20b (AMD platforms use ROCm's FP8/quantized path; native MXFP4 acceleration is exclusive to NVIDIA Blackwell machines), with headroom left for concurrency and long context — ready out of the box.
AI-Medium (AMD Radeon AI PRO R9700 32GB GPU, preloaded with gpt-oss-20b)
Get quote →
32B-class high-quality inference (coding assistance / advanced Q&A)
約 27 萬
Qwen3 32B · Gemma 3 27B
18–24GB (4-bit); Q8: Gemma 27B ≈27–30GB (fits a single 32GB card), Qwen 32B ≈34–35GB (slightly over 32GB)
64GB DDR5・16-core Ryzen 9 9950X3D
At 4-bit, a single 32GB card runs it with room for context; Gemma 27B fits even at Q8, while Qwen 32B at Q8 slightly exceeds 32GB and needs a shorter sequence or 48GB. The difference from AI-Medium (NT$151K / Radeon AI PRO R9700 · ROCm): NVIDIA's CUDA ecosystem, long-run-stable Pro drivers, and Gemma preloaded.
AI-Medium-Gemma (RTX PRO 4500 32GB)
Get quote →

Large LLM Local Inference (70B–120B)

70B-class models need roughly 43GB VRAM at 4-bit quantization, putting a single 48GB professional card within range; MAQ's AI-High uses an RTX PRO 5000 48GB card, leaving about 5GB of headroom for 4-bit 70B — enough for typical context lengths. The comparable Qwen3 72B at Q4 needs about 47GB, pushing close to the limit, so long context requires dropping to a lower quantization (IQ4_XS). Full precision, the 120B MoE model, or comfortably running long context requires a single 96GB card, or even a multi-GPU WRX90 platform doing tensor parallelism. We recommend 128GB+ system memory, ECC preferred, to support long-context KV cache and CPU offload.

Use case · representative models Recommended spec Key differences Recommended MAQ build
Local 70B inference (4-bit, single-card entry point)
Llama 3.3 70B (Qwen3 72B-class comparable)
Llama 70B ≈43GB (Q4_K_M, including KV cache); 72B-class Q4 ≈47GB, near the limit on a 48GB card
128GB DDR5 ECC
24-core Threadripper 9960X
A single RTX PRO 5000 48GB card runs 4-bit Llama 70B (≈43GB) with about 5GB of headroom — enough for typical context lengths; the 72B-class model at Q4 (≈47GB) nearly fills VRAM, so long context needs a lower quantization such as IQ4_XS. For comfortable long context, full precision, or high-concurrency service, move up to a single 96GB card or multiple GPUs.
AI-High (RTX PRO 5000 48GB)
約 61 萬
Get quote →
Single 96GB card running gpt-oss-120b, no model splitting needed
gpt-oss-120b (MoE, MXFP4) / high-precision 70B
gpt-oss-120b weights ≈60GB (116.8B total params / 5.1B active), a single 96GB card leaves ample room for KV cache
256GB DDR5 ECC
32-core Threadripper PRO 9975WX
A single RTX PRO 6000 96GB card fully carries gpt-oss-120b — no need to split across GPUs; the same 96GB card can also run near-full-precision 70B, giving the simplest deployment and lowest latency.
AI-Highend (RTX PRO 6000 96GB)
約 120 萬
Get quote →
Multi-GPU tensor parallelism: full-precision 70B or distributed 120B
Full-precision Llama 70B (FP16 ≈140GB) / distributed gpt-oss-120b
≥140GB combined across GPUs; the WRX90 platform natively supports 4–7 pooled 96GB cards
256GB DDR5 ECC (expandable further)
96-core Threadripper PRO 9995WX
Full-precision 70B (FP16 ≈140GB) does not fit on a single card and requires cross-GPU tensor parallelism; WRX90's native 4–7-card configuration suits full-precision 70B or high-concurrency 120B service, where inter-GPU bandwidth is the key factor.
AMD-WRX90 (4–7 GPU multi-card platform)
約 152 萬
Get quote →
Local 70B inference (4-bit, single-card entry point)
約 61 萬
Llama 3.3 70B (Qwen3 72B-class comparable)
Llama 70B ≈43GB (Q4_K_M, including KV cache); 72B-class Q4 ≈47GB, near the limit on a 48GB card
128GB DDR5 ECC・24-core Threadripper 9960X
A single RTX PRO 5000 48GB card runs 4-bit Llama 70B (≈43GB) with about 5GB of headroom — enough for typical context lengths; the 72B-class model at Q4 (≈47GB) nearly fills VRAM, so long context needs a lower quantization such as IQ4_XS. For comfortable long context, full precision, or high-concurrency service, move up to a single 96GB card or multiple GPUs.
AI-High (RTX PRO 5000 48GB)
Get quote →
Single 96GB card running gpt-oss-120b, no model splitting needed
約 120 萬
gpt-oss-120b (MoE, MXFP4) / high-precision 70B
gpt-oss-120b weights ≈60GB (116.8B total params / 5.1B active), a single 96GB card leaves ample room for KV cache
256GB DDR5 ECC・32-core Threadripper PRO 9975WX
A single RTX PRO 6000 96GB card fully carries gpt-oss-120b — no need to split across GPUs; the same 96GB card can also run near-full-precision 70B, giving the simplest deployment and lowest latency.
AI-Highend (RTX PRO 6000 96GB)
Get quote →
Multi-GPU tensor parallelism: full-precision 70B or distributed 120B
約 152 萬
Full-precision Llama 70B (FP16 ≈140GB) / distributed gpt-oss-120b
≥140GB combined across GPUs; the WRX90 platform natively supports 4–7 pooled 96GB cards
256GB DDR5 ECC (expandable further)・96-core Threadripper PRO 9995WX
Full-precision 70B (FP16 ≈140GB) does not fit on a single card and requires cross-GPU tensor parallelism; WRX90's native 4–7-card configuration suits full-precision 70B or high-concurrency 120B service, where inter-GPU bandwidth is the key factor.
AMD-WRX90 (4–7 GPU multi-card platform)
Get quote →

LLM Fine-Tuning (LoRA / QLoRA)

Fine-tuning is bound by VRAM and system memory, not raw compute. QLoRA compresses the base model to 4-bit and trains only a small LoRA adapter, lowering the 8B–70B barrier to what a single card can handle; but training runs continuously for hours, and a silent memory soft error can corrupt an entire run, so 70B-class fine-tuning is best paired with ECC memory and generous system RAM (to leave room for the optimizer / CPU offload). Figures below are conservative estimates for QLoRA 4-bit, batch size 1, sequence length 2048, with gradient checkpointing enabled.

Use case · representative models Recommended spec Key differences Recommended MAQ build
Entry-level fine-tuning: 8B model customization (support tone / domain Q&A)
Llama 3.1 8B · Qwen3 8B · Gemma 3 12B
≈14–16GB (QLoRA 4-bit); a 24GB card allows longer context
32GB DDR5 minimum
12-core Ryzen 9
8B QLoRA is the most comfortable sweet spot for a single 24GB card; for larger datasets or CPU offload, we recommend bumping system memory to 64GB.
AI-Eco (RTX PRO 4000 24GB)
約 10 萬
Get quote →
Mid-size fine-tuning: 13B–32B (multi-turn instruction tuning / style alignment)
Qwen3 14B · Gemma 3 27B · Qwen3 32B (tight fit)
14B ≈20–22GB; 27B ≈24–28GB; 32B ≈28–32GB (QLoRA 4-bit + checkpointing)
64GB DDR5
16-core Ryzen 9 9950X3D
27B QLoRA pushes close to 32GB once the optimizer is running, so a 32GB card is needed for a usable batch size / context; 32B only works with short sequences on 32GB — for long sequences, we recommend 48GB (AI-High).
AI-Medium-Gemma (RTX PRO 4500 32GB)
約 27 萬
Get quote →
Local 70B flagship fine-tuning (QLoRA, single card)
Llama 3.3 70B · Qwen3 72B-class
≈46–48GB (QLoRA 4-bit, ≈7K sequence length); standard LoRA (unquantized) ≈160GB
256GB DDR5 ECC
32-core Threadripper PRO 9975WX
70B QLoRA needs roughly 46–48GB. It fits on a 48GB card, but sequence length and batch size have to run right at the ceiling; a 96GB card leaves room to extend sequences, increase batch size, and hold optimizer state. ECC plus 256GB RAM protects long training runs from being ruined by a soft error. For full-precision LoRA (≈160GB), see the multi-GPU platform.
AI-Highend (RTX PRO 6000 96GB)
約 120 萬
Get quote →
Full-precision LoRA / multi-GPU distributed 70B fine-tuning (research-grade)
Llama 3.3 70B · Qwen3 72B (bf16 LoRA)
Standard LoRA ≈160GB → requires multiple GPUs (e.g. 2×96GB); a single 96GB card can run high-precision QLoRA + long sequences
256GB DDR5 ECC
32–96-core Threadripper PRO
WRX90's native 4–7-card setup suits unquantized 70B LoRA or running multiple fine-tuning experiments concurrently; if a single 96GB card is enough for high-precision QLoRA, consider AI-Highend instead (approximately NT$1,200,000).
AMD-WRX90 (4–7 GPU multi-card platform)
約 152 萬
Get quote →
Entry-level fine-tuning: 8B model customization (support tone / domain Q&A)
約 10 萬
Llama 3.1 8B · Qwen3 8B · Gemma 3 12B
≈14–16GB (QLoRA 4-bit); a 24GB card allows longer context
32GB DDR5 minimum・12-core Ryzen 9
8B QLoRA is the most comfortable sweet spot for a single 24GB card; for larger datasets or CPU offload, we recommend bumping system memory to 64GB.
AI-Eco (RTX PRO 4000 24GB)
Get quote →
Mid-size fine-tuning: 13B–32B (multi-turn instruction tuning / style alignment)
約 27 萬
Qwen3 14B · Gemma 3 27B · Qwen3 32B (tight fit)
14B ≈20–22GB; 27B ≈24–28GB; 32B ≈28–32GB (QLoRA 4-bit + checkpointing)
64GB DDR5・16-core Ryzen 9 9950X3D
27B QLoRA pushes close to 32GB once the optimizer is running, so a 32GB card is needed for a usable batch size / context; 32B only works with short sequences on 32GB — for long sequences, we recommend 48GB (AI-High).
AI-Medium-Gemma (RTX PRO 4500 32GB)
Get quote →
Local 70B flagship fine-tuning (QLoRA, single card)
約 120 萬
Llama 3.3 70B · Qwen3 72B-class
≈46–48GB (QLoRA 4-bit, ≈7K sequence length); standard LoRA (unquantized) ≈160GB
256GB DDR5 ECC・32-core Threadripper PRO 9975WX
70B QLoRA needs roughly 46–48GB. It fits on a 48GB card, but sequence length and batch size have to run right at the ceiling; a 96GB card leaves room to extend sequences, increase batch size, and hold optimizer state. ECC plus 256GB RAM protects long training runs from being ruined by a soft error. For full-precision LoRA (≈160GB), see the multi-GPU platform.
AI-Highend (RTX PRO 6000 96GB)
Get quote →
Full-precision LoRA / multi-GPU distributed 70B fine-tuning (research-grade)
約 152 萬
Llama 3.3 70B · Qwen3 72B (bf16 LoRA)
Standard LoRA ≈160GB → requires multiple GPUs (e.g. 2×96GB); a single 96GB card can run high-precision QLoRA + long sequences
256GB DDR5 ECC・32–96-core Threadripper PRO
WRX90's native 4–7-card setup suits unquantized 70B LoRA or running multiple fine-tuning experiments concurrently; if a single 96GB card is enough for high-precision QLoRA, consider AI-Highend instead (approximately NT$1,200,000).
AMD-WRX90 (4–7 GPU multi-card platform)
Get quote →

Image Generation (Stable Diffusion / Flux / ComfyUI)

What drives VRAM usage in image generation is not model size but what has to stay resident at once — the diffusion model itself, the VAE, and the text encoder (Flux and SD3.5 both carry T5-XXL) all need to be on the card the moment an image is generated. SDXL runs in 8–12GB once quantized, but Flux.1 dev (12B) needs 24GB or more for full precision; high-quality Flux LoRA training alone needs 48GB, and training and inference resident together requires 96GB. As of 2026, ComfyUI remains the dominant local workflow. VRAM figures below are conservative ranges with the quantization stated.

Use case · representative models Recommended spec Key differences Recommended MAQ build
SDXL entry to advanced (ComfyUI workflows / style LoRA inference)
SDXL 1.0 · Illustrious XL (fp16)
8–12GB (fp16, 1024²); stacking two ControlNet + IP-Adapter sets adds another 5–8GB
32GB DDR5
12-core Ryzen 9
24GB is more than enough for pure SDXL inference and can even handle multiple stacked ControlNets; it can also just about fit Flux at fp8. A reasonable starting point on the tightest budget.
AI-Eco (RTX PRO 4000 24GB)
約 10 萬
Get quote →
Main ComfyUI rig: mixed Flux.1 + SDXL, batch generation
Flux.1 dev/schnell (12B) · SD3.5 Large (8B) · SDXL
Flux.1 fp8 ≈17GB, Q4 GGUF ≈12GB; SD3.5 Large fp8 ≈11–14GB
64GB DDR5
16-core Ryzen 9 9950X3D
The RTX PRO 4500 32GB runs Flux.1 fp8 and SD3.5 Large smoothly, making it the 32GB-tier pick for mixed Flux generation; for full fp16 Flux.1 (24GB+), step up to the 48GB AI-High; for training LoRA and running inference simultaneously, you need the 96GB single-card AI-Highend. NVIDIA CUDA ecosystem, long-run-stable Pro drivers.
AI-Medium-Gemma (RTX PRO 4500 32GB)
約 27 萬
Get quote →
Simultaneous LoRA training and inference / full-precision Flux (commercial production line)
Flux.1 dev (fp16 + LoRA training) · SD3.5 Large · SDXL
48GB recommended for Flux LoRA training; running training and inference together needs 48GB for headroom (conservative estimate)
256GB DDR5 ECC
32-core Threadripper PRO 9975WX
High-quality Flux LoRA training alone needs about 48GB; to also reserve inference VRAM alongside training without hitting OOM, the threshold for "both resident at once" lands on a single 96GB card. The 48GB AI-High can still do it, but training and inference must be time-sliced. ECC memory suits an uninterrupted commercial production line.
AI-Highend (RTX PRO 6000 96GB)
約 120 萬
Get quote →
SDXL entry to advanced (ComfyUI workflows / style LoRA inference)
約 10 萬
SDXL 1.0 · Illustrious XL (fp16)
8–12GB (fp16, 1024²); stacking two ControlNet + IP-Adapter sets adds another 5–8GB
32GB DDR5・12-core Ryzen 9
24GB is more than enough for pure SDXL inference and can even handle multiple stacked ControlNets; it can also just about fit Flux at fp8. A reasonable starting point on the tightest budget.
AI-Eco (RTX PRO 4000 24GB)
Get quote →
Main ComfyUI rig: mixed Flux.1 + SDXL, batch generation
約 27 萬
Flux.1 dev/schnell (12B) · SD3.5 Large (8B) · SDXL
Flux.1 fp8 ≈17GB, Q4 GGUF ≈12GB; SD3.5 Large fp8 ≈11–14GB
64GB DDR5・16-core Ryzen 9 9950X3D
The RTX PRO 4500 32GB runs Flux.1 fp8 and SD3.5 Large smoothly, making it the 32GB-tier pick for mixed Flux generation; for full fp16 Flux.1 (24GB+), step up to the 48GB AI-High; for training LoRA and running inference simultaneously, you need the 96GB single-card AI-Highend. NVIDIA CUDA ecosystem, long-run-stable Pro drivers.
AI-Medium-Gemma (RTX PRO 4500 32GB)
Get quote →
Simultaneous LoRA training and inference / full-precision Flux (commercial production line)
約 120 萬
Flux.1 dev (fp16 + LoRA training) · SD3.5 Large · SDXL
48GB recommended for Flux LoRA training; running training and inference together needs 48GB for headroom (conservative estimate)
256GB DDR5 ECC・32-core Threadripper PRO 9975WX
High-quality Flux LoRA training alone needs about 48GB; to also reserve inference VRAM alongside training without hitting OOM, the threshold for "both resident at once" lands on a single 96GB card. The 48GB AI-High can still do it, but training and inference must be time-sliced. ECC memory suits an uninterrupted commercial production line.
AI-Highend (RTX PRO 6000 96GB)
Get quote →

AI Agent / Agentic Workflows (n8n / LangGraph / CrewAI)

This workload's bottleneck is CPU core count, memory, and network I/O — not a high-end discrete GPU. Orchestration engines mostly run a loop of "call the LLM API → wait for the response → parse JSON → move to the next node," with each tool call occupying a CPU thread; running many agents in parallel is really about core count and 32GB of memory carrying the load. VRAM only starts to matter once you move inference back on-premise (for privacy, offline use, or cutting API costs) — and agentic workflows mostly only need lightweight 8B–20B local models, not 70B.

Use case · representative models Recommended spec Key differences Recommended MAQ build
Pure orchestration: multi-agent flows, all LLM calls via cloud API
Cloud API-based (local machine only handles embeddings / classification)
0GB discrete GPU requirement (integrated graphics suffice); local embeddings ≈1–2GB
32GB DDR5
6–10 cores (core count matters more than clock speed)
Orchestration is I/O- and thread-intensive, not compute-intensive; 6–10 cores + 32GB is enough to run dozens of agents in parallel. Preloaded with n8n / LangGraph / CrewAI / Ollama / Open WebUI, ready out of the box — spending on a high-end discrete GPU here would be wasted.
AI-Agent-Medium / Eco
約 3.3 萬
Get quote →
Semi-local: some nodes switched to a lightweight local LLM (8B class)
Qwen3 8B · Llama 3.1 8B (Q4_K_M)
≈5–8GB (4-bit, within 16K context)
32GB DDR5
6–10 cores
An 8B-class model quantized down lands at 5–8GB, which even integrated graphics or an entry-level iGPU can drive at roughly 30–40 tok/s — enough for classification, extraction, and tool routing within an agent flow. A discrete GPU is only needed for higher speed or higher concurrency.
AI-Agent-Medium / Eco
約 3.3 萬
Get quote →
Local inference core: primary LLM moved on-premise (20B class, offline / no API cost)
gpt-oss-20b (MXFP4) · Qwen3 14B
≈16GB (gpt-oss-20b is MoE + MXFP4, weights ≈12–13GB, ≈16GB at runtime); 24–32GB recommended for long context
64GB DDR5
12 cores
Upgrade to this tier once API bills start to hurt, or when compliance requires data to stay on-premise. Both AI-Medium (Radeon AI PRO R9700 32GB) and AI-Eco (RTX PRO 4000 24GB) leave context headroom, and come preloaded with gpt-oss-20b / Llama, ready to plug straight into your orchestrator.
AI-Medium / AI-Eco
約 16 萬
Get quote →
Heavy local agent: 70B-class complex reasoning chains / long-document agents
Llama 3.3 70B · Qwen3 72B (Q4_K_M)
≈42–45GB (4-bit); 48GB is the safe single-card floor, which is exactly what this machine provides
128GB DDR5 ECC
24-core Threadripper
Most agentic workflows never need 70B — this tier is for cases requiring fully autonomous, high-quality local reasoning chains. For full precision or longer context, look to a dual-GPU or 96GB machine.
AI-High (RTX PRO 5000 48GB)
約 61 萬
Get quote →
Pure orchestration: multi-agent flows, all LLM calls via cloud API
約 3.3 萬
Cloud API-based (local machine only handles embeddings / classification)
0GB discrete GPU requirement (integrated graphics suffice); local embeddings ≈1–2GB
32GB DDR5・6–10 cores (core count matters more than clock speed)
Orchestration is I/O- and thread-intensive, not compute-intensive; 6–10 cores + 32GB is enough to run dozens of agents in parallel. Preloaded with n8n / LangGraph / CrewAI / Ollama / Open WebUI, ready out of the box — spending on a high-end discrete GPU here would be wasted.
AI-Agent-Medium / Eco
Get quote →
Semi-local: some nodes switched to a lightweight local LLM (8B class)
約 3.3 萬
Qwen3 8B · Llama 3.1 8B (Q4_K_M)
≈5–8GB (4-bit, within 16K context)
32GB DDR5・6–10 cores
An 8B-class model quantized down lands at 5–8GB, which even integrated graphics or an entry-level iGPU can drive at roughly 30–40 tok/s — enough for classification, extraction, and tool routing within an agent flow. A discrete GPU is only needed for higher speed or higher concurrency.
AI-Agent-Medium / Eco
Get quote →
Local inference core: primary LLM moved on-premise (20B class, offline / no API cost)
約 16 萬
gpt-oss-20b (MXFP4) · Qwen3 14B
≈16GB (gpt-oss-20b is MoE + MXFP4, weights ≈12–13GB, ≈16GB at runtime); 24–32GB recommended for long context
64GB DDR5・12 cores
Upgrade to this tier once API bills start to hurt, or when compliance requires data to stay on-premise. Both AI-Medium (Radeon AI PRO R9700 32GB) and AI-Eco (RTX PRO 4000 24GB) leave context headroom, and come preloaded with gpt-oss-20b / Llama, ready to plug straight into your orchestrator.
AI-Medium / AI-Eco
Get quote →
Heavy local agent: 70B-class complex reasoning chains / long-document agents
約 61 萬
Llama 3.3 70B · Qwen3 72B (Q4_K_M)
≈42–45GB (4-bit); 48GB is the safe single-card floor, which is exactly what this machine provides
128GB DDR5 ECC・24-core Threadripper
Most agentic workflows never need 70B — this tier is for cases requiring fully autonomous, high-quality local reasoning chains. For full precision or longer context, look to a dual-GPU or 96GB machine.
AI-High (RTX PRO 5000 48GB)
Get quote →

Desktop Mini AI Development Machine (GB10 Unified Memory)

This tier makes a different trade-off from GPU-based workstations: instead of "one large graphics card + system memory," it uses an NVIDIA GB10 Grace Blackwell Superchip where CPU and GPU share a single 128GB pool of unified memory. The upside is loading 100B-class models for local inference and development on one machine, in a case about the size of a mini PC, with no server room or dedicated power required; the trade-off is memory bandwidth (LPDDR5X, approximately 273GB/s) far below data-center HBM, so token throughput lands in the tens per second. It is positioned for "local development, validation, and always-on agents," not high-concurrency inference serving. Well suited to individuals or small teams bringing AI development to their desk, with data never leaving the company.

Use case · representative models Recommended spec Key differences Recommended MAQ build
Desk-side local development: single machine running 100B-class model inference
Nemotron 3 Super 120B (MoE) · Llama 3.3 70B · Qwen3 80B (quantized)
Shared 128GB unified memory (CPU+GPU pool); a 120B-class model at 4-bit quantization takes ≈60GB
128GB unified memory
20-core ARM (Cortex-X925 + A725)
Uses the same GB10 Superchip and 128GB unified memory as the NVIDIA DGX Spark, differing in its 1TB SSD and brand, at a lower price. A lower-cost entry point to the GB10 platform for local LLM development and inference; generation speed is bounded by memory bandwidth, and it is not built for high-concurrency serving.
ASUS Ascent GX10 (GB10 / 1TB)
約 18 萬
Get quote →
Always-on agents / a dev machine needing larger local storage
Same as above, also suited to long-running autonomous agents
Shared 128GB unified memory
128GB unified memory
20-core ARM (Cortex-X925 + A725)
The NVIDIA-branded DGX Spark, with a 4TB NVMe SSD for more local storage, suited to developers who need to keep large numbers of models and datasets on the machine itself. Its unified memory architecture matches the GX10; the reason to choose it is mainly the larger storage and NVIDIA-branded provenance.
NVIDIA DGX Spark (GB10 / 4TB)
約 21 萬
Get quote →
Desk-side local development: single machine running 100B-class model inference
約 18 萬
Nemotron 3 Super 120B (MoE) · Llama 3.3 70B · Qwen3 80B (quantized)
Shared 128GB unified memory (CPU+GPU pool); a 120B-class model at 4-bit quantization takes ≈60GB
128GB unified memory・20-core ARM (Cortex-X925 + A725)
Uses the same GB10 Superchip and 128GB unified memory as the NVIDIA DGX Spark, differing in its 1TB SSD and brand, at a lower price. A lower-cost entry point to the GB10 platform for local LLM development and inference; generation speed is bounded by memory bandwidth, and it is not built for high-concurrency serving.
ASUS Ascent GX10 (GB10 / 1TB)
Get quote →
Always-on agents / a dev machine needing larger local storage
約 21 萬
Same as above, also suited to long-running autonomous agents
Shared 128GB unified memory
128GB unified memory・20-core ARM (Cortex-X925 + A725)
The NVIDIA-branded DGX Spark, with a 4TB NVMe SSD for more local storage, suited to developers who need to keep large numbers of models and datasets on the machine itself. Its unified memory architecture matches the GX10; the reason to choose it is mainly the larger storage and NVIDIA-branded provenance.
NVIDIA DGX Spark (GB10 / 4TB)
Get quote →
Why not DIY it yourself

Enterprise AI adoption is hard — MAQ handles it before the machine ships

Picking the right spec is only step one. What really costs an IT team time is the environment setup and tuning pitfalls after the bare metal arrives — MAQ does that work before it ever leaves the factory.

Buying bare metal from a major brand yourself

  • A bare machine arrives with no environment — your IT team spends two weeks or more just on Ubuntu, CUDA drivers, and Docker permissions.
  • Misconfigured PCIe lane allocation can cut multi-GPU performance in half, and the cause is hard to trace.
  • Specs get bought on a hunch — overspend on excess capacity, or get stuck short — with no one to run the numbers for you.

MAQ: ready out of the box, tuned end to end

  • Ships with Ollama (Qwen / Gemma), ComfyUI, and an enterprise knowledge base preinstalled — plug in power and network, and you're serving an API within 5 minutes.
  • Engineers run stress testing and PCIe topology tuning before shipping, so multi-GPU PCIe 5.0 lanes are correctly allocated and performance is maxed out.
  • A specialist builds the optimal GPU configuration based on your budget and model size — that's exactly the comparison table you're looking at above.
End-to-end hardware + software service

Environment set up, plug-and-play, backed by local support

From a single phone call to a live API on boot, MAQ covers all of it. You never have to touch a driver, chase down PCIe lanes, or run your own stress tests — just focus on your AI application.

1

Spec consultation

Tell us your budget and the models you'll run — a specialist builds an optimized GPU configuration, with no overspend and no bottlenecks.

2

Custom build

Industrial-grade components, enterprise-class power and cooling, and multi-GPU PCIe 5.0 topology optimization.

3

Stress testing + environment setup

Burn-in stress testing, with Ollama / ComfyUI / CUDA / Docker / your enterprise knowledge base preinstalled.

4

Engineer-delivered go-live

Personal delivery and on-site inspection anywhere in Taiwan, including outlying islands — serving an API within 5 minutes of power-on.

5

Local technical support

3-year hardware warranty, loaner units for contracted accounts, and expert-level Proxmox VE / hardware troubleshooting support, remote and on-site.

On-premise generative AI for enterprise · Hardware + software, unified

MAQ Alishan Enterprise AI Knowledge Appliance

Turn documents, SOPs, and database knowledge into a private AI you can simply ask. RAG knowledge base × knowledge graph × permission auditing all stay on your own AI server — confidential data never leaves the premises. Two tiers, matched to scale:

Alishan Standard

SMBs · Departmental knowledge base
  • A single 32–48GB-class GPU, running 7B–32B models locally
  • Internal RAG knowledge base setup (HR / IT / sales Q&A)
  • Access control and audit logging, data stays on your LAN
Learn about Standard →
Flagship

Alishan Flagship

Large research labs · Data centers
  • RTX PRO 6000 96GB or multiple Blackwell GPUs in tandem
  • Load the full gpt-oss-120b on a single card, or fine-tune Llama 3.3 70B
  • High concurrency, knowledge graphs, enterprise-grade auditing — measured at 161 tok/s
Learn about Flagship →
Our enterprise service commitment

Buy MAQ, and you're never left to figure out AI alone

From pre-purchase spec consultation to post-deployment technical support and warranty coverage, our Taiwan-based team has had you covered since 2002.

Custom hardware spec consultation

Based on your budget and the model sizes you'll run, a specialist builds an optimized GPU / memory / CPU configuration — no overspend, no bottlenecks.

On-premise data security

All computation and data stay 100% on-premise. Models, knowledge bases, and conversation logs never leave your LAN — no risk of personal or trade-secret data exposure.

Local warranty and technical support in Taiwan

Proxmox VE virtualization environment, expert-level hardware troubleshooting, a 3-year hardware warranty, loaner units for contracted accounts, and engineer delivery with on-site inspection anywhere in Taiwan, including outlying islands.

FAQ

Frequently asked AI purchasing questions

How should you configure hardware to run Llama 3.3 70B locally in 2026? How much VRAM do you need?

At 4-bit (Q4_K_M) quantization you need roughly 43GB VRAM, which a single 48GB professional card can handle, with 128GB ECC system memory recommended. The comparable Qwen3 72B at Q4 needs about 47GB, pushing close to the limit on a 48GB card, so long context requires a lower quantization (IQ4_XS); the matching MAQ build, AI-High, uses an RTX PRO 5000 48GB card and leaves about 5GB of headroom running 4-bit 70B. For comfortable long context, full precision, or high concurrency, move up to a single 96GB card or multiple GPUs. The matching MAQ model is AI-High (approximately NT$610,000).

Do you need a multi-GPU server to deploy gpt-oss-120b? Can a single 96GB card handle it?

A single card is enough — no multi-GPU setup required. gpt-oss-120b uses a MoE architecture with native MXFP4 quantization, with weights around 60GB; a single NVIDIA RTX PRO 6000 Blackwell 96GB card fully loads it with plenty of room left for KV cache, giving the simplest deployment and lowest latency. Only full precision or high-throughput, high-concurrency serving requires the multi-GPU WRX90 platform. The matching MAQ model is AI-Highend (approximately NT$1,200,000).

For a Flux.1 dev image generation workstation, how much VRAM and memory avoids OOM errors?

When generating with Flux.1 dev (12B), the diffusion model itself, the VAE, and the T5-XXL text encoder must all stay resident on the card at once: fp8 needs about 17GB, full fp16 about 24GB; SDXL fp16 needs about 8–12GB. A 32GB card (such as the RTX PRO 4500) runs Flux.1 fp8 and SD3.5 smoothly; for training LoRA and running inference together, 48GB gives more comfortable headroom. We recommend 64GB system memory as a starting point. The matching MAQ model for image generation is AI-Medium-Gemma (RTX PRO 4500 32GB, approximately NT$270,000).

Do I need a high-end graphics card to set up an AI Agent workflow with n8n or LangGraph?

In most cases, no. n8n / LangGraph / CrewAI mainly call cloud or lightweight local LLM APIs, so the bottleneck is CPU core count and 32GB of memory, not a high-end discrete GPU. A GPU is only needed once inference moves fully on-premise (for privacy, offline use, or cutting API costs), and even then a lightweight 8B–20B model is usually enough. The MAQ AI Agent PC comes preloaded with n8n / LangGraph / CrewAI / Ollama, at approximately NT$34,000.

How do you configure multiple GPUs? How many cards do you need for full-precision 70B or distributed 120B?

Full-precision Llama 70B (FP16, approximately 140GB) does not fit on a single card and requires multi-GPU tensor parallelism. The AMD WRX90 platform natively supports 4–7 GPUs pooled together, suited to full-precision 70B or distributed high-concurrency 120B serving, where inter-GPU bandwidth is the key factor. The matching MAQ model is AMD-WRX90 (4–7 pooled RTX PRO 6000 96GB cards, approximately NT$1,520,000).

How much VRAM and memory do you need to fine-tune Llama 70B locally?

QLoRA 4-bit fine-tuning of 70B needs roughly 46–48GB VRAM, putting a single 48GB card at the entry spec; the AI-High (RTX PRO 5000 48GB) fits, but has to run right at the ceiling, with conservative sequence length and batch size — so the recommended model here is the 96GB single-card AI-Highend. We recommend ECC memory plus 128GB+ system RAM to protect long training runs from being ruined by a memory soft error. Standard (unquantized) LoRA needs roughly 160GB, requiring multiple GPUs or a large 96GB card. The matching MAQ model is AI-High (approximately NT$610,000).

For running a 100B-class model on a single desk machine, should you choose a mini machine like DGX Spark or a GPU workstation?

It depends on your use case. The NVIDIA DGX Spark and ASUS Ascent GX10 use the NVIDIA GB10 Superchip with 128GB of unified memory, letting a single machine load 100B-class models for local inference and development, in a case about the size of a mini PC and needing no server room — well suited to local development, validation, and always-on agents. But their memory bandwidth (approximately 273GB/s) is below data-center-class hardware, so token throughput lands in the tens per second, making them unsuitable for high-concurrency external services. For high-throughput inference serving multiple users, or for image generation and large-scale fine-tuning, a GPU-based workstation (such as AI-High / AI-Highend) is a better fit. Matching MAQ models: ASUS Ascent GX10 (approximately NT$180,000), NVIDIA DGX Spark (approximately NT$210,000).

Still not sure what to configure?

Tell MAQ the model, use case, and budget — our engineers will size it exactly right, no overspend, no bottlenecks.

See the Alishan on-premise enterprise solution → · Further reading: MAQ blog buying guides and comparisons →