Aug 15, 2026 β€’ ai-research

Pestle-27B-Ternary Review 2026: A 27B Model in an 8.5 GB GGUF for Local Medical & General Use

Pestle-27B-Ternary compresses a 27B model to a single 8.48 GB GGUF via ternary weights, targeting local medical QA, biomedical retrieval, coding, and general assistance. We review the benchmarks and trade-offs.

Running a 27B-class model locally used to mean 50+ GB of weights. Pestle-27B-Ternary packages comparable capability into a single 8.48 GB GGUF.

What is Pestle-27B-Ternary?

Pestle-27B-Ternary is a compact 27B ternary-weight language model for local inference, built on Qwen3.6-27B with Doses AI’s ternary compression (weights constrained to -1/0/+1) plus a matching-parent BF16 final decoder block. It bundles private medical QA, biomedical evidence work, pharmaceutical retrieval, coding, and general assistance into one runnable GGUF, served by the Mortar runtime (llama.cpp-compatible). It’s explicitly a research preview, not a medical device.

Key features

  • 27B-class model compressed to a single 8.48 GB GGUF via ternary weights (-1/0/+1)
  • Strong medical benchmarks: MedQA 89.79, MedMCQA 68.85, PubMedQA 76.70 accuracy
  • Runs locally with Mortar on Apple Silicon, NVIDIA CUDA, or CPU fallback
  • General capability retained: MMLU-Redux 83.53, GSM8K 93.25, HumanEval+ 87.20
  • Up to 262K context; optional vision input via a separate mmproj projection file

Who should use it?

Researchers, clinicians-in-training, and developers building local, privacy-preserving medical-text or biomedical-retrieval assistants who want 27B-class quality without server-grade VRAM. It’s also a capable general/coding model for anyone who can run an 8.5 GB GGUF.

Pros and cons

Pros: dramatic size compression with strong medical scores; fully local and Apache-2.0; broad general capability.

Cons: research preview only (not for clinical use); requires building/running the separate Mortar runtime; compression trades some accuracy versus full-precision FP16.

Pricing

Free open weights under the Apache-2.0 license.

FAQ

Can I run it on a Mac? Yes β€” Mortar uses Metal on Apple Silicon; CPU-only builds also work, just slower.

Is it a medical device? No β€” it’s a research preview and must not be used for diagnosis or treatment decisions.

What base model does it use? Qwen3.6-27B, with ternary compression applied across the language model.

Explore the best AI Research & Alignment tools

Related Articles

Subscribe to the 9bests weekly β€” get the full list free

Hand-picked AI tool reviews and updates every week. Subscribe to receive this full list + 7 more quick-reference sheets (writing / image / video / audio / chat models / data / API cost).

Subscribe free & get it β†’

Independent reviews β€” ratings aren't influenced by vendor payments Β· double opt-in Β· unsubscribe anytime