Browse

Agens Volundr 32B Preview

Updated October 6, 2026

Overview / Description

Agens Volundr 32B Preview is an open-source AI large language model built for long-context coding, tool use, and extended agent sessions on self-hosted hardware. It is a dense 32-billion-parameter model (no mixture-of-experts) released under the Apache-2.0 license, with a 262K-token context window and a memory design that keeps the resource footprint predictable. Instead of letting the KV cache grow without limit, 54 of its 72 layers use a fixed-size state, and it adds Blockway Compressed-Sparse Attention for selective long-range token access plus an Engram n-gram memory that lives in host RAM rather than GPU memory. The practical result is that it runs on modest hardware: the BF16 build fits on two 48GB GPUs or a single H200, while INT4 quantization drops it to 31.7 GiB and runs on one 48GB GPU, sustaining about 24 tokens/second decode from 8K up to 128K context. For teams that want a commercial-grade model they can run in-house without cloud dependencies, this is a concrete option, and weights are available now on Hugging Face. BestAIFor readers evaluating self-hosted LLMs will find the fixed-state layer approach the main differentiator to test against the context length they actually need.

Used For

Self-hosting an open-source LLM for long-context coding, tool use, and extended agent sessions without cloud dependencies

Pricing

Plan

Free

Free — Apache-2.0 open-source model, download weights from Hugging Face; self-hosting hardware costs apply

View pricing

Pros & Cons

Pros

  • 262K-token context window with a fixed-size state across 54 of 72 layers to cap memory growth
  • INT4 quantization fits 31.7 GiB on a single 48GB GPU; BF16 runs on two 48GB GPUs or one H200
  • Apache-2.0 licensed, so commercial self-hosting is permitted with weights on Hugging Face
  • Dense (non-MoE) design gives a predictable, flat resource footprint
  • Engram n-gram memory stored in host RAM keeps GPU memory free for the context

Cons

  • Labeled a 32B 'Preview', so stability and benchmarks may still be in flux
  • Still needs 48GB-class GPUs even at INT4, out of reach for consumer cards
  • 24 tokens/second decode is modest compared with hosted frontier APIs
  • Self-hosting requires in-house ML-ops skills to deploy and maintain

Questions & Answers

Alternatives

Qwen2.5-32B, Mistral Small, and DeepSeek-V2

Reviews & Ratings

—

0 reviews

5
0%
4
0%
3
0%
2
0%
1
0%

Sign in to rate and review Agens Volundr 32B Preview.

Sign in to review

No reviews yet. Be the first to review Agens Volundr 32B Preview!

Try Agens Volundr 32B Preview free