Pestle-27B-Ternary Review 2026: A 27B Model in an 8.5 GB GGUF for Local Medical & General Use
Pestle-27B-Ternary compresses a 27B model to a single 8.48 GB GGUF via ternary weights, targeting local medical QA, biomedical retrieval, coding, and general assistance. We review the benchmarks and trade-offs.
Running a 27B-class model locally used to mean 50+ GB of weights. Pestle-27B-Ternary packages comparable capability into a single 8.48 GB GGUF.
What is Pestle-27B-Ternary?
Pestle-27B-Ternary is a compact 27B ternary-weight language model for local inference, built on Qwen3.6-27B with Doses AIβs ternary compression (weights constrained to -1/0/+1) plus a matching-parent BF16 final decoder block. It bundles private medical QA, biomedical evidence work, pharmaceutical retrieval, coding, and general assistance into one runnable GGUF, served by the Mortar runtime (llama.cpp-compatible). Itβs explicitly a research preview, not a medical device.
Key features
- 27B-class model compressed to a single 8.48 GB GGUF via ternary weights (-1/0/+1)
- Strong medical benchmarks: MedQA 89.79, MedMCQA 68.85, PubMedQA 76.70 accuracy
- Runs locally with Mortar on Apple Silicon, NVIDIA CUDA, or CPU fallback
- General capability retained: MMLU-Redux 83.53, GSM8K 93.25, HumanEval+ 87.20
- Up to 262K context; optional vision input via a separate mmproj projection file
Who should use it?
Researchers, clinicians-in-training, and developers building local, privacy-preserving medical-text or biomedical-retrieval assistants who want 27B-class quality without server-grade VRAM. Itβs also a capable general/coding model for anyone who can run an 8.5 GB GGUF.
Pros and cons
Pros: dramatic size compression with strong medical scores; fully local and Apache-2.0; broad general capability.
Cons: research preview only (not for clinical use); requires building/running the separate Mortar runtime; compression trades some accuracy versus full-precision FP16.
Pricing
Free open weights under the Apache-2.0 license.
FAQ
Can I run it on a Mac? Yes β Mortar uses Metal on Apple Silicon; CPU-only builds also work, just slower.
Is it a medical device? No β itβs a research preview and must not be used for diagnosis or treatment decisions.
What base model does it use? Qwen3.6-27B, with ternary compression applied across the language model.
Explore the best AI Research & Alignment tools
Related Articles
Why AI Agent Memory Should Decay: A Hands-On Test of AIOBR
Remembering more does not always make an AI agent smarter. We tested how AIOBR uses world versioning, decay, trajectories, skill compression, and counterfactual learning to govern long-term memory.
ModelMap Review 2026: AI Benchmarks as a 3D Spikiness Map
A review of ModelMap (modelmap.tech) β an interactive 3D visualization that turns AI model benchmark scores into explorable shapes, parsed live from Hugging Face model cards.
Open Science Desktop Review 2026: A Local-First AI Research Workbench
Open Science Desktop is a local-first, model-agnostic AI research workbench that runs the whole research loop β survey, experiment, analysis, write-up β in one auditable session. We review what it does and who it's for.
RLHF and AI Alignment in 2026: From Rules to Character
The latest breakthroughs in AI alignment β how the field moved from hand-crafted rules to training AI systems with stable behavioral traits that generalize across domains.
Subscribe to the 9bests weekly β get the full list free
Hand-picked AI tool reviews and updates every week. Subscribe to receive this full list + 7 more quick-reference sheets (writing / image / video / audio / chat models / data / API cost).
Subscribe free & get it βIndependent reviews β ratings aren't influenced by vendor payments Β· double opt-in Β· unsubscribe anytime