Deploying Qwen on Intel Arc with vLLM and DFlash2

I've been running Qwen3.6-35B-A3B with INT4 weights on a single Intel Arc Pro B70 with 32 GB of GPU memory, using vLLM's XPU backend. A DFlash2 draft model adds speculative decoding, and an OpenAI-compatible API makes the deployment usable from ordinary chat clients.

The interesting part is fitting the target model, draft model, and a useful context budget into the same GPU while keeping generation responsive. This post describes the deployment snapshot recorded on October 6, 2026, along with retained benchmark results. The measurements were not rerun for this post. Addresses and filesystem paths in the examples are generic replacements.

The serving stack

The request path is straightforward:

Qwen inference architecture: an API client connects to vLLM, which coordinates the INT4 target and DFlash2 draft model on a shared Intel Arc Pro B70 GPU.

vLLM owns inference and scheduling. Clients supply messages and, optionally, tool definitions. The server can return structured tool calls, but the client application is responsible for executing them and returning their results.

The recorded runtime consists of:

Component Deployment snapshot
Operating system Ubuntu 26.04 LTS, x86_64
CPU Intel Core i7-12700K
System memory Approximately 61 GiB visible to Linux
GPU Intel Arc Pro B70, 32 GB
Python 3.12.14
vLLM 0.29.0
PyTorch 2.13.0+xpu
XPU kernels vllm-xpu-kernels 0.1.14.1, locally modified
Transformers 5.17.0
compressed-tensors 0.17.0

The launcher initializes the Intel runtime and selects the Level Zero GPU backend with ONEAPI_DEVICE_SELECTOR=level_zero:gpu. System RAM and swap do not extend the GPU's allocation budget: weights, caches, and execution buffers still have to fit on the device.

INT4 weights are only one part of the memory budget

The target is the amd/Qwen3.6-35B-A3B-w4a16-llmcompressor checkpoint. The deployment records describe GPTQ weight-only INT4 quantization in compressed-tensors format, using symmetric groups of 128 and BF16 activations.

Different parts of the runtime use different precisions:

Allocation Selected precision
Target weights Packed INT4
General computation BF16
Target attention KV cache auto, identified as BF16 in the selected benchmark
Recurrent / Mamba state cache FP16
Draft attention KV cache FP8

Quantizing the weights does not make every allocation INT4. The target attention cache alone receives a fixed 6.5 GiB allocation, while the draft model, recurrent state, runtime buffers, and target weights consume additional memory.

This distinction matters when increasing context length. Earlier experiments with FP32 recurrent state and automatic sizing failed at the full configured context; the deployed recurrent cache remains FP16.

Context length, scheduling, and caching

The server is configured for a 262,144-token context window, including input and generated output. That is a configured ceiling; the retained long-prompt benchmark used 133,745 input tokens.

Setting Selected value
Maximum model length 262,144 tokens
Maximum active sequences 1
Batched token limit 8,192
Attention block size 64
Target KV allocation 6.5 GiB
Prefix caching Enabled
Async scheduling Enabled
XPU graphs Enabled by default
Ahead-of-time compilation Disabled
Reasoning parser qwen3
Automatic tool choice parser qwen3_coder

The 8,192-token scheduler budget is separate from the total context limit. Likewise, async scheduling does not create more active sequence slots: with one active sequence, competing requests can queue.

Prefix caching helps when requests reuse eligible prefixes. Stable instructions and shared document prefixes can reduce repeated prompt processing, but a changed prefix, cache eviction, or server restart can remove that benefit.

DFlash2 speculative decoding

The draft model proposes tokens for the target model to verify. Whether this pays off depends on the cost of drafting and how many proposals the target accepts.

The selected configuration uses three speculative tokens per step and an FP8 draft KV cache. With a generic local model path, its --speculative-config value is:

{
  "method": "dflash",
  "model": "/opt/models/qwen-dflash2",
  "num_speculative_tokens": 3,
  "kv_cache_dtype": "fp8"
}

The installed speculative implementation requires min_p=0 and does not support logit_bias, according to the experiment record. These are constraints of this deployment, not a promise about every vLLM release.

The custom launcher also supports omitting speculative decoding for a target-only diagnostic run. That requires restarting the server with a different launch configuration; it is not a per-request toggle. Keeping that mode available helps isolate draft-model compatibility and memory issues.

A minimal API check

Use a short served-model name rather than exposing a local checkpoint path to clients. This example uses qwen-int4 as a placeholder; replace it with the ID returned by /v1/models on your own server.

curl --fail --silent --show-error \
  http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "qwen-int4",
    "messages": [{"role": "user", "content": "Reply with exactly READY."}],
    "temperature": 0,
    "min_p": 0,
    "max_tokens": 16,
    "chat_template_kwargs": {"enable_thinking": false}
  }'

A warm loopback request with these generation settings returned READY. in approximately 0.116 seconds, using 17 prompt tokens and 3 completion tokens. This is a short readiness check, not a throughput benchmark.

The reasoning parser does not force every request to produce reasoning. In this check, the request explicitly disables thinking through chat_template_kwargs.

For routine inspection, /health checks readiness, /v1/models confirms the served identity, and /metrics exposes request and cache behavior. I look at running requests, waiting requests, and KV cache utilization together: an available API can still have a queue.

Long-prompt performance

All of the following recorded runs used 133,745 prompt tokens. They compare complete configurations and decoding conditions; they are not controlled measurements of one isolated optimization.

Recorded run Output tokens Prefill tokens/s Decode tokens/s Time to first token
DFlash2, 3 draft tokens, FP8 draft KV; cold greedy 512 3,785 112.05 35.49 s
Same selected setup; cached greedy 512 — 107.14 0.97 s
Same selected setup; sampled 105 3,788 78.11 35.51 s
Final setup with fixed 6.5 GiB target KV; cold greedy 512 3,797 106.67 35.40 s

The selected experiment peaked at approximately 30,368 MiB of GPU memory. The final fixed-cache run recorded approximately 30,171 MiB. Both are close enough to the device's capacity that additional allocations deserve careful measurement.

A decode rate above 100 tokens per second can coexist with a 35-second wait before the first token. The server has to process the long prompt before generating an answer. The cached run shows how prefix reuse can change that wait substantially without increasing the decode rate.

For interactive use, I would measure first-token latency and queue time alongside decode throughput. Trimming unnecessary history and preserving reusable prefixes can matter more than chasing a higher peak generation rate.

These results have limits. Sampled and greedy runs produced different output lengths and text. Only output-coherence spot checks were recorded, and there was no full quality evaluation. A separate run affected by GPU resets was excluded. The measurements do not establish full-context stability or sustained performance under mixed-client load.

Local patches are part of the deployment

Package versions alone do not reproduce this environment. Two local changes are important:

  • The DFlash implementation was patched so quantized context projections use their projection modules rather than assuming directly accessible dense weights. Unquantized projections retain the fused path.
  • The XPU kernels binary was replaced. Deployment notes attribute it to a stacked-weight RMSNorm fix, but the documentation inspection verified the binary difference without independently establishing its build provenance.

This makes the setup a record of a working environment, rather than an installation recipe guaranteed to work from stock packages. Before rebuilding it, preserve the model revision, launcher, dependency versions, source patches, and the provenance of any custom binaries.

An upgrade should be checked with model startup, a short chat, structured tool output, and a representative long request before it replaces the working environment.

Keeping the service predictable

The deployment uses a user systemd unit with restart-on-failure behavior and a controlled stop timeout. The unit and launcher prevent an alternate backend from starting alongside it, since both would compete for the same GPU and port.

Unattended startup needs its own check. An enabled user unit does not by itself prove that the service starts at boot without a login. User lingering was disabled in the recorded snapshot, so boot availability remains unverified.

A restart interrupts active requests, clears runtime caches, and requires model initialization again. Environment changes belong in the persisted launcher or unit configuration; changing an interactive shell does not update a running service.

For a deployment example, loopback is a useful default API boundary. Remote access should be configured deliberately through an authenticated gateway or private network, with transport protection appropriate to that route. API compatibility alone does not provide an access boundary.

The remaining checks are full-context quality and stability, sustained mixed-client load, unattended boot, and image-input validation. Text generation was verified; multimodal behavior was not tested in this snapshot.

The main lesson from this setup is to budget the entire GPU workload and measure the stages separately. Weight quantization makes the model fit, cache precision shapes the context budget, speculative decoding can improve generation, and prefix reuse can reduce the wait before generation starts. Each contributes to the result, and each needs its own evidence.