Crawl4AI
Crawl4AI is an open-source and LLM-friendly web crawler and scraper that facilitates efficient data extraction from the web.
This category lists open-source data science tools for integration, visualization, and analytics. Airbyte moves data from APIs, databases, and files into warehouses and lakes via ETL/ELT pipelines. Apache Superset covers interactive data exploration and charts, and ClickHouse is a real-time analytics database for large volumes of data.
41 tools available
Crawl4AI is an open-source and LLM-friendly web crawler and scraper that facilitates efficient data extraction from the web.
A comprehensive collection of algorithms and data structures implemented in JavaScript, complete with explanations and additional reading links.
PandasAI allows you to interact with your data sources such as databases and datalakes using conversational language. It leverages Language Model Models (LLMs) and RAG to provide intuitive data analysis.
Robust Speech Recognition via Large-Scale Weak Supervision, capable of transcribing and translating spoken language.
Apache Superset is an open-source data visualization and data exploration platform, offering a wide range of visualization options and a user-friendly interface for interactive data exploration.
A curated list of awesome Machine Learning frameworks, libraries, and software, providing insights into various useful resources in the domain.
A collection of extracted system prompts from popular chatbots, including ChatGPT, Claude, and Gemini.
Metabase is an easy-to-use open source Business Intelligence and Embedded Analytics tool that enables seamless interaction with data for everyone.
MindsDB is an AI's query engine, a platform for building AI models that can learn and answer questions over large-scale federated data.
pdfplumber is a powerful Python library for extracting detailed information from PDF files, including text, tables, individual characters, shapes, and more. It allows precise programmatic access to underlying PDF content, making it highly useful for data extraction, automation, and analysis tasks.
The Prometheus monitoring system and time series database for collecting and storing metrics as time series data.
Lightdash is a self-serve business intelligence (BI) tool designed to empower data teams to work more efficiently and effectively.
Our Newsletter
Get short emails with useful data science projects, releases, and repos worth watching.
Graphiti helps to build real-time knowledge graphs tailored for AI agents, enhancing their decision-making and data processing capabilities.
State-of-the-art Machine Learning library for Pytorch, TensorFlow, and JAX, providing thousands of pre-trained models for natural language processing, computer vision, and other areas.
The leading data integration platform for ETL / ELT data pipelines from APIs, databases, and files to data warehouses, data lakes, and data lakehouses. Both self-hosted and Cloud-hosted.
A minimalist deep learning framework designed to be simple and easy to understand, inspired by pytorch and micrograd.
STORM is an LLM-powered knowledge curation system designed to research topics and generate comprehensive full-length reports with citations. Developed by Stanford's OVAL team, STORM leverages large language models to streamline information gathering and synthesis.
Cocoindex is a powerful data transformation framework designed for AI applications. It is ultra performant, offering incremental processing capabilities.
Specification and comprehensive documentation for the Model Context Protocol, a standard for managing AI model contextual data efficiently.
A high-performance, cloud-native vector database designed for scalable and efficient vector Approximate Nearest Neighbor (ANN) search.
AI orchestration framework to build customizable, production-ready LLM applications. Connect components such as models, vector databases, and file converters to pipelines or agents that can interact with your data. Best suited for building Retrieval-Augmented Generation (RAG), question answering, semantic search, or conversational agent chatbots.
Laminar - an open-source all-in-one platform for engineering AI products. Create a data flywheel for your AI applications, with features such as Traces, Evals, Datasets, and Labels. Part of the Y Combinator Summer 2024 batch.
Cube is a universal semantic layer platform for AI, BI, spreadsheets, and embedded analytics. It serves as an access layer for data analytics, providing real-time data access and integration across multiple sources.
A curated and ranked collection of top machine learning libraries and tools for Python. Regularly updated to highlight the best open-source machine learning projects in the Python ecosystem.
Make Your Company Data Driven. Connect to any data source, easily visualize, dashboard, and share your data.
Handsontable is a JavaScript data grid and data table offering a spreadsheet-like look and feel, fully compatible with React, Angular, and Vue. Developed and supported by the Handsontable team.
Open-source IoT Platform for comprehensive device management, data collection, processing, and visualization.
Argilla is a collaborative platform designed for AI engineers and domain experts to efficiently create, curate, and manage high-quality datasets for AI and machine learning projects.
Jitsu is an open-source Segment alternative with a fully-scriptable data ingestion engine designed for modern data teams. Set up a real-time data pipeline in minutes, not days.
Querybook is a Big Data Querying UI that combines collocated table metadata with a simple notebook interface for efficient data analysis.
Get notified about new tools and updates to existing ones.