Skip to content

PaperQA

FutureHouse technical staff / FutureHouse

PaperQA provides agentic retrieval-augmented question answering over scientific documents, combining local search, ranked evidence summaries and answers with in-text citations.

Catalog updated ·

Overview

PaperQA, described in the project README as PaperQA2, is a Python package for retrieval-augmented generation focused on scientific literature. It accepts PDFs, text, Microsoft Office documents and source code files, placing it in the document-search and evidence-synthesis stage of a research workflow rather than serving as a scientific predictive model. Users can interact through the pqa command-line interface or integrate its Python API into their own applications.

The documented agent workflow searches a local document index, chunks and embeds candidate documents, retrieves relevant passages, and uses language models to score and summarize evidence in relation to a question. Selected summaries then support answer generation with in-text citations. The agent can repeat or reorder these operations to refine its search. Python responses expose the answer, formatted answer, question and evidence context; users who want explicit control over document selection can instead add and query documents through Docs without agent orchestration.

PaperQA also retrieves paper metadata from external providers, including Crossref and Semantic Scholar, and reuses local indexes on subsequent queries. Language models and embeddings are configurable through LiteLLM-based interfaces, with documented options for locally hosted models and alternative vector stores. These dependencies are separate from the repository itself. An LLM service or local server is required, and the README cautions that the public package does not include all internal FutureHouse paper-access tools. Users must supply their own PDFs, so published research results should not be assumed to describe performance on an arbitrary local collection.

Key Features

  • Agentic search, evidence gathering and answer generation, with tools that can be repeated or invoked in different orders.
  • Passage retrieval followed by LLM-based relevance scoring and contextual summarization to support answers with in-text citations.
  • Local full-text indexing with reuse on later queries and synchronization when documents are added.
  • Automatic paper metadata retrieval from multiple providers, including citation information and retraction checks described in the quickstart.
  • Document ingestion for PDFs, text, Microsoft Office documents and source code, with documented readers for figures and tables.
  • Configurable LLMs, embeddings and vector stores, plus synchronous and asynchronous Python interfaces and manual `Docs` workflows.

Use Cases

  • Suggested evaluation: answer a focused chemistry literature question from a curated local paper collection, then check each cited passage against the original document.
  • Suggested evaluation: use the documented `contracrow` configuration to investigate whether supplied papers contain evidence contradicting a specific scientific claim.
  • Suggested evaluation: integrate the Python API into a literature-review workflow that retains both generated answers and the supporting evidence summaries for researcher inspection.

How to Use

  1. Read the repository README to choose between the CLI, agentic Python API and manual Docs workflow. The supplied package metadata requires Python 3.11 or later.
  2. Follow the README installation guidance for the paper-qa package. Check optional reader dependencies if your collection includes Office documents or PDF media.
  3. Configure an LLM provider and embeddings using the documented settings. Consult the linked LiteLLM provider documentation for provider requirements, or follow the README's local-server examples.
  4. Assemble a small, relevant collection of papers and select its directory. Use the documented pqa question-answering workflow or Python ask interface to build the index and retrieve evidence.
  5. Inspect the formatted answer, citations and evidence context against the source documents. For an intended evaluation, repeat representative questions with controlled settings and record unsupported claims or missed evidence before expanding the collection.

Related resources

ChemCrow is a Python chemistry agent package built with Langchain that connects language-model reasoning to RDKit, paper-qa, chemical databases, and reaction-planning tools.

Open sourcePython

Cheminformatics · Literature & Research

ChemDataExtractor v2 is a Python toolkit for extracting chemical entities, properties, spectra and tabulated data from scientific literature, with readers for HTML, XML and PDF.

Open sourcePython

Literature & Research

Coscientist provides a simple implementation and supporting research data for large-language-model chemical research, covering synthesis planning, library selections, repeated runs and optimization.

Literature & Research · Lab Automation

PubMed MCP (cyanheads) is a third-party MCP server connecting AI clients to biomedical literature search, article metadata, available full text, MeSH vocabulary and citation tools.

Open sourceTypeScript

Literature & Research · Drug Discovery

SciAgents is a research multi-agent framework that uses scientific knowledge graphs, LLMs and retrieval tools to develop and critique hypotheses for bio-inspired materials research.

Open sourcePython

Literature & Research · Materials Discovery

STELLA

Agent

STELLA is a research multi-agent framework for biomedical literature and bioinformatics workflows, combining reusable workflow templates, tool orchestration, critique, and optional dynamic tool creation.

Open sourcePython

Literature & Research · Drug Discovery

Related guides

Agents

Agents and Skills for Scientific Workflows

Plan bounded scientific agent workflows with inspectable inputs, limited tools, reusable skills, evidence-linked outputs and explicit human approval. Compare the roles of ChemCrow, PaperQA, Coscientist and Scientific Agent Skills without assuming tested interoperability.