Pestle-27B-Ternary
Pestle-27B-Ternary is a compact 27B ternary-weight language model (8.48 GB GGUF) for local inference, packing private medical QA, biomedical evidence, pharmaceutical retrieval, coding, and general assistance into one runnable file under the Mortar runtime — a research preview, not a medical device.
✅ Pros / Advantages
- • 27B-class model compressed to a single 8.48 GB GGUF via ternary weights (-1/0/+1)
- • Strong medical benchmarks: MedQA 89.79, MedMCQA 68.85, PubMedQA 76.70 accuracy
- • Runs locally with Mortar (llama.cpp-compatible) on Apple Silicon, NVIDIA CUDA, or CPU
- • General capability retained: MMLU-Redux 83.53, GSM8K 93.25, HumanEval+ 87.20
- • Up to 262K context; optional vision input via a separate mmproj projection file
❌ Cons / Limitations
- • Research preview only — explicitly not for clinical/diagnostic use
- • Requires building/running the separate Mortar runtime (no one-click hosted endpoint)
- • Based on Qwen3.6-27B; compression trades some accuracy vs full-precision FP16
💰 Pricing Plans
Free (Open Weights, Apache-2.0)
Pricing details are gathered from public sources and are subject to change. Please visit the official website for real-time rates.
Last updated: July 2026 · 9bests editorial review
✅ Who should use Pestle-27B-Ternary
- • 27B-class model compressed to a single 8.48 GB GGUF via ternary weights (-1/0/+1)
- • Strong medical benchmarks: MedQA 89.79, MedMCQA 68.85, PubMedQA 76.70 accuracy
- • Runs locally with Mortar (llama.cpp-compatible) on Apple Silicon, NVIDIA CUDA, or CPU
- • General capability retained: MMLU-Redux 83.53, GSM8K 93.25, HumanEval+ 87.20
- • Up to 262K context; optional vision input via a separate mmproj projection file
⚠️ Who should look elsewhere
- • Research preview only — explicitly not for clinical/diagnostic use
- • Requires building/running the separate Mortar runtime (no one-click hosted endpoint)
- • Based on Qwen3.6-27B; compression trades some accuracy vs full-precision FP16
🎯 Common use cases
Literature and paper review
Experiment tracking
Synthesis and alignment research
⚖️ Pestle-27B-Ternary vs SkillSpector
| Pestle-27B-Ternary | SkillSpector | |
|---|---|---|
| Rating | 4.35/5 | 4.3/5 |
| Pricing | Free (Open Weights, Apache-2.0) | Free (Open Source) |
| Key strength | 27B-class model compressed to a single 8.48 GB GGUF via ternary weights (-1/0/+1) | Boosts workflow efficiency |
See the full head-to-head in our Pestle-27B-Ternary vs SkillSpector comparison.
❓ Frequently asked questions
Is Pestle-27B-Ternary free?
+
Pestle-27B-Ternary offers a free tier (Free (Open Weights, Apache-2.0)). Paid plans unlock higher limits and advanced features.
What is Pestle-27B-Ternary best for?
+
Pestle-27B-Ternary is best for 27B-class model compressed to a single 8.48 GB GGUF via ternary weights (-1/0/+1) and Strong medical benchmarks: MedQA 89.79, MedMCQA 68.85, PubMedQA 76.70 accuracy. Pestle-27B-Ternary is a compact 27B ternary-weight language model (8.48 GB GGUF) for local inference, packing private medical QA, biomedical evidence, pharmaceutical retrieval, coding, and general assistance into one runnable file under the Mortar runtime — a research preview, not a medical device.
How does Pestle-27B-Ternary compare to SkillSpector?
+
Pestle-27B-Ternary (4.35/5) and SkillSpector (4.3/5) serve overlapping needs. Pestle-27B-Ternary stands out for 27B-class model compressed to a single 8.48 GB GGUF via ternary weights (-1/0/+1), while SkillSpector is stronger at Boosts workflow efficiency. Choose based on your priority.
🔄 Top Alternatives to Pestle-27B-Ternary
Related ToolsSkillSpector
NVIDIA-developed AI agent skill security scanner that detects vulnerabilities, malicious patterns, and safety risks in AI agent skills.
Memento
Self-hosted agentic search and LLM wiki over your email archive, turning decades of emails into a personal wiki with People, Projects, Concepts, and Newsletters dimensions.
Reyn
Always-on local-first AI that watches your screen, journals your work, and provides instant search across everything you've worked on.
ModelMap
ModelMap (modelmap.tech) is an interactive 3D visualization that turns AI model benchmark scores into explorable shapes. Each model's performance across public benchmarks is rendered as a 'spiky' 3D form — longer spikes mean higher scores — parsed live from Hugging Face model cards. Built on an open-source '3D Graph' library, it offers a flight-simulator-style interface (WASD to fly, mouse to look, click a spike to zoom, hover for tooltips) for browsing model data in space rather than static tables. A hidden Star Wars-themed mini-game underscores its goal of making model analysis more playful. It's a free, browser-based research toy — novel for building intuition, though it has drawn technical criticism on how it represents scores.