Cataloged from orcarouter/Qwen3.8-27B-Uncensored-MLX
An abliterated (refusal-removed) build of
Qwen/Qwen3.8-27Bβ a 27B-parameter dense, hybrid-attention (Gated DeltaNet linear + full attention) native vision-language model with thinking control, tool-calling and an MTP head β quantized to MLX format for Apple Silicon. Four precisions are provided β 2 / 4 / 6 / 8-bit (affine, group size 64) β each as a subfolder, with the 4-bit build also mirrored at the repo root so thatorcarouter/Qwen3.8-27B-Uncensored-MLXloads directly in LM Studio and other tools that treat a repo as a single model. The vision tower, norms and conv layers are kept in BF16; only the language-model linear weights (includingembed_tokens/lm_head) are quantized. Browse all models in the OrcaRouter Model Catalog. This model is deployed as API here.
This model has had its safety alignment substantially removed via abliteration (orthogonalizing the refusal direction out of the residual stream). As a direct consequence:
Qwen3.8-27B would refuse. It has no meaningful built-in guardrails.By downloading or using this model you acknowledge and accept the above.
| Folder | Bits/weight | Size | Shards | Min Mac RAM | Quality vs BF16 source |
|---|---|---|---|---|---|
8-bit/ | 8.627 | ~27.5 GB | 6 | 32 GB | Near-lossless β recommended for quality |
6-bit/ | 6.661 | ~22 GB | 5 | 24β32 GB | Excellent β strong quality/size balance |
4-bit/ | 4.695 | ~15 GB | 3 | 24 GB | Very good β recommended default |
2-bit/ | 2.729 | ~8.7 GB | 2 | 16 GB | β οΈ Severely degraded β archival only |
2-bit warning: at 27B, 2-bit quantization collapses generation quality (repetition loops, garbled output). It is included only as an extreme-compression archive; do not use it for real work β prefer 4-bit or higher.
Repo root =
4-bit/. The root of this repo holds a copy of the 4-bit build, so--model orcarouter/Qwen3.8-27B-Uncensored-MLX(no subfolder) resolves to 4-bit. Use the subfolder paths to pick any other precision.
All builds were quantized from the same abliterated BF16 source and verified numerically (dequantized weights vs. source) plus tested by generation on GPU.
| Precision | Numerical fidelity (cosine) | Text / Chinese / Code | Refusal probes | Vision |
|---|---|---|---|---|
| 8-bit | cos 0.9997 | β | β 0 refusals | β |
| 6-bit | cos 0.9996 | β | β 0 refusals | β |
| 4-bit | cos 0.996 | β | β 0 refusals | β |
| 2-bit | cos 0.92 | β οΈ breaks down | β οΈ garbled (not refusal) | partial |
Note: on 6-bit, mlx's offline
mx.dequantizemis-unpacks these weights (a library edge case), so correctness is verified by clean generation β inference is unaffected.
pip install -U mlx-vlm # needs mlx-vlm >= 0.6.13, mlx >= 0.32
# download one precision (e.g. 4-bit) from the subfolder
hf download orcarouter/Qwen3.8-27B-Uncensored-MLX --include "4-bit/*" \
--local-dir ./Qwen3.8-27B-Uncensored-MLX
# text
python -m mlx_vlm generate \
--model ./Qwen3.8-27B-Uncensored-MLX/4-bit \
--prompt "Explain quantum entanglement in one sentence." --max-tokens 256
# vision (image + text)
python -m mlx_vlm generate \
--model ./Qwen3.8-27B-Uncensored-MLX/4-bit \
--image path/to/image.png \
--prompt "Describe this image." --max-tokens 256
# OpenAI-compatible server
python -m mlx_vlm server --model ./Qwen3.8-27B-Uncensored-MLX/4-bit --port 8080
On Apple Silicon the Metal backend is used automatically β no CUDA setup needed.
(On a Linux CUDA backend, vision requires MLX_CUDA_USE_CUDNN_SDPA=0; this does not
apply on macOS.)
This model has a native MTP head. In MLX, MTP is loaded as a separate drafter for
speculative decoding: the main model is loaded with the MTP weights stripped, and the drafter
is passed explicitly. The drafter lives in the mtp/ subfolder of this repo
(model_type: qwen3_5_mtp) and works with any main-model precision (4 / 6 / 8-bit).
Setting an
mtp_enabledflag on the main model alone does nothing β MLX needs the separate drafter passed via--draft-model β¦ --draft-kind mtp.
# fetch a main-model precision (e.g. 6-bit) plus the MTP drafter
hf download orcarouter/Qwen3.8-27B-Uncensored-MLX --include "6-bit/*" "mtp/*" \
--local-dir ./Qwen3.8-27B-Uncensored-MLX
# generate with MTP speculative decoding
python -m mlx_vlm generate \
--model ./Qwen3.8-27B-Uncensored-MLX/6-bit \
--draft-model ./Qwen3.8-27B-Uncensored-MLX/mtp \
--draft-kind mtp --draft-block-size 4 \
--prompt "Explain quantum entanglement in one sentence." --max-tokens 256
# OpenAI-compatible server with MTP
python -m mlx_vlm server \
--model ./Qwen3.8-27B-Uncensored-MLX/6-bit \
--draft-model ./Qwen3.8-27B-Uncensored-MLX/mtp \
--draft-kind mtp --draft-block-size 4 --port 8080
Requirements: an mlx-vlm build with the qwen3_5_mtp drafter and --draft-kind mtp
(available on mlx-vlm main). MTP acceptance is lossless β with greedy decoding the output is
identical to running without the drafter, just fewer forward passes on accepted tokens. The
speedup is realized on Apple Silicon (Metal); one drafter serves all precisions.
Search for orcarouter/Qwen3.8-27B-Uncensored-MLX in LM Studio and download it β the repo
root is the 4-bit build, and the other precisions appear as separate download options.
Three things to get right:
If you are on an older LM Studio MLX runtime, update it (Settings β Runtime): qwen3_5
support landed in mlx-vlm 0.6.x, and older runtimes cannot load this architecture at all.
| Base model | Qwen/Qwen3.8-27B |
| Architecture | Qwen3_5ForConditionalGeneration β 64 layers, hidden 5120, hybrid Gated DeltaNet (48 linear + 16 full attention, interval 4), native VL tower |
| Modification | Abliteration (refusal-direction removal), then MLX affine quantization |
| Quantization | MLX affine, group size 64, per-precision 2 / 4 / 6 / 8-bit |
| Kept in BF16 | vision tower, all norms, linear-attention conv1d |
| Quantized | language-model linear layers incl. embed_tokens and lm_head |
| Context | 262,144 tokens |