Overview / Description
Agens Volundr 32B Preview is an open-source AI large language model built for long-context coding, tool use, and extended agent sessions on self-hosted hardware. It is a dense 32-billion-parameter model (no mixture-of-experts) released under the Apache-2.0 license, with a 262K-token context window and a memory design that keeps the resource footprint predictable. Instead of letting the KV cache grow without limit, 54 of its 72 layers use a fixed-size state, and it adds Blockway Compressed-Sparse Attention for selective long-range token access plus an Engram n-gram memory that lives in host RAM rather than GPU memory. The practical result is that it runs on modest hardware: the BF16 build fits on two 48GB GPUs or a single H200, while INT4 quantization drops it to 31.7 GiB and runs on one 48GB GPU, sustaining about 24 tokens/second decode from 8K up to 128K context. For teams that want a commercial-grade model they can run in-house without cloud dependencies, this is a concrete option, and weights are available now on Hugging Face. BestAIFor readers evaluating self-hosted LLMs will find the fixed-state layer approach the main differentiator to test against the context length they actually need.
Used For
Self-hosting an open-source LLM for long-context coding, tool use, and extended agent sessions without cloud dependencies
Pricing
Plan
Free — Apache-2.0 open-source model, download weights from Hugging Face; self-hosting hardware costs apply
Pros & Cons
Pros
- 262K-token context window with a fixed-size state across 54 of 72 layers to cap memory growth
- INT4 quantization fits 31.7 GiB on a single 48GB GPU; BF16 runs on two 48GB GPUs or one H200
- Apache-2.0 licensed, so commercial self-hosting is permitted with weights on Hugging Face
- Dense (non-MoE) design gives a predictable, flat resource footprint
- Engram n-gram memory stored in host RAM keeps GPU memory free for the context
Cons
- Labeled a 32B 'Preview', so stability and benchmarks may still be in flux
- Still needs 48GB-class GPUs even at INT4, out of reach for consumer cards
- 24 tokens/second decode is modest compared with hosted frontier APIs
- Self-hosting requires in-house ML-ops skills to deploy and maintain
Questions & Answers
Alternatives
Qwen2.5-32B, Mistral Small, and DeepSeek-V2
Reviews & Ratings
0 reviews
Sign in to rate and review Agens Volundr 32B Preview.
Sign in to reviewNo reviews yet. Be the first to review Agens Volundr 32B Preview!