AI alignment, safety research, and model development insights
Self-hosted agentic search and LLM wiki over your email archive, turning decades of emails into a personal wiki with People, Projects, Concepts, and Newsletters dimensions.
A local-first, model-agnostic AI research workbench for macOS, Windows & Linux. It runs the whole research loop β exploration, literature survey, hypothesis, experiment code, analysis, figures, and write-up β in one auditable, reproducible desktop session.
NVIDIA-developed AI agent skill security scanner that detects vulnerabilities, malicious patterns, and safety risks in AI agent skills.
Always-on local-first AI that watches your screen, journals your work, and provides instant search across everything you've worked on.
A CLI from Experiential Labs that turns collected agent traces into smaller models you own. wmo optimize distills frontier behaviour via the Tinker API, and wmo serve routes requests between frontier and small models β reported at frontier-level quality for 27% less cost on RouterBench.
ModelMap (modelmap.tech) is an interactive 3D visualization that turns AI model benchmark scores into explorable shapes. Each model's performance across public benchmarks is rendered as a 'spiky' 3D form β longer spikes mean higher scores β parsed live from Hugging Face model cards. Built on an open-source '3D Graph' library, it offers a flight-simulator-style interface (WASD to fly, mouse to look, click a spike to zoom, hover for tooltips) for browsing model data in space rather than static tables. A hidden Star Wars-themed mini-game underscores its goal of making model analysis more playful. It's a free, browser-based research toy β novel for building intuition, though it has drawn technical criticism on how it represents scores.
Our #1 pick is Memento (4.5/5), followed by Open Science Desktop (4.5/5). Rankings are based on hands-on testing across features, pricing, and real-world performance.
Every tool is scored on a 5-point scale across features, ease of use, pricing & value, reliability, and integrations. We re-test monthly so rankings stay current.