Tools for web scraping, data extraction, and AI data pipelines
Open-source web crawler designed for LLMs and AI agents with structured extraction and browser automation.
Open-source local-first cognitive memory system implementing AGM-compatible belief revision that automatically re-evaluates downstream beliefs when facts change, with SHA-256 hash chain for data integrity.
A deterministic SQL semantic inspector that catches silently-wrong AI-generated queries β double-counting, bad joins, exposed PII β in about 0.1 ms before they run. Works as a CI gate, an MCP server, or a library.
Adaptive Recall is a hosted memory system for AI applications that goes far beyond simple vector search. It stores, recalls, and manages long-term memory for agents and apps over MCP or a plain REST API, and β unlike a static embeddings store β it actively learns. Four retrieval strategies run in parallel (vector similarity, temporal recency, full-text keyword, and knowledge-graph traversal), and the system learns which to prioritize for each query type. Results are ranked with ACT-R cognitive scoring from 30 years of cognitive-science research, factoring in recency, access frequency, entity connections, and validated confidence. A knowledge graph is built automatically from stored memories, memories move through a confidence-based lifecycle and fade when unused, and an ML pipeline trains on your usage patterns β validating every parameter change against real query history before adopting it. A simple eight-tool API (store, recall, update, forget, graph, status, snapshot, feedback) covers everything, with Bearer-token auth and JSON in/out. Free, Starter, Pro, and Business plans are available.
ParseHawk is a fully local document AI processing toolkit β no data leaves your machine. It ships with an API server, CLI, and Web UI, making it easy to integrate into existing workflows or use standalone for document parsing, chunking, OCR, and Q&A over documents.
Our #1 pick is Crawl4AI (4.3/5), followed by Atlas (4.3/5). Rankings are based on hands-on testing across features, pricing, and real-world performance.
Every tool is scored on a 5-point scale across features, ease of use, pricing & value, reliability, and integrations. We re-test monthly so rankings stay current.