Overview
PaperQA, described in the project README as PaperQA2, is a Python package for retrieval-augmented generation focused on scientific literature. It accepts PDFs, text, Microsoft Office documents and source code files, placing it in the document-search and evidence-synthesis stage of a research workflow rather than serving as a scientific predictive model. Users can interact through the pqa command-line interface or integrate its Python API into their own applications.
The documented agent workflow searches a local document index, chunks and embeds candidate documents, retrieves relevant passages, and uses language models to score and summarize evidence in relation to a question. Selected summaries then support answer generation with in-text citations. The agent can repeat or reorder these operations to refine its search. Python responses expose the answer, formatted answer, question and evidence context; users who want explicit control over document selection can instead add and query documents through Docs without agent orchestration.
PaperQA also retrieves paper metadata from external providers, including Crossref and Semantic Scholar, and reuses local indexes on subsequent queries. Language models and embeddings are configurable through LiteLLM-based interfaces, with documented options for locally hosted models and alternative vector stores. These dependencies are separate from the repository itself. An LLM service or local server is required, and the README cautions that the public package does not include all internal FutureHouse paper-access tools. Users must supply their own PDFs, so published research results should not be assumed to describe performance on an arbitrary local collection.
Key Features
- Agentic search, evidence gathering and answer generation, with tools that can be repeated or invoked in different orders.
- Passage retrieval followed by LLM-based relevance scoring and contextual summarization to support answers with in-text citations.
- Local full-text indexing with reuse on later queries and synchronization when documents are added.
- Automatic paper metadata retrieval from multiple providers, including citation information and retraction checks described in the quickstart.
- Document ingestion for PDFs, text, Microsoft Office documents and source code, with documented readers for figures and tables.
- Configurable LLMs, embeddings and vector stores, plus synchronous and asynchronous Python interfaces and manual `Docs` workflows.
Use Cases
- Suggested evaluation: answer a focused chemistry literature question from a curated local paper collection, then check each cited passage against the original document.
- Suggested evaluation: use the documented `contracrow` configuration to investigate whether supplied papers contain evidence contradicting a specific scientific claim.
- Suggested evaluation: integrate the Python API into a literature-review workflow that retains both generated answers and the supporting evidence summaries for researcher inspection.
How to Use
- Read the repository README to choose between the CLI, agentic Python API and manual
Docsworkflow. The supplied package metadata requires Python 3.11 or later. - Follow the README installation guidance for the paper-qa package. Check optional reader dependencies if your collection includes Office documents or PDF media.
- Configure an LLM provider and embeddings using the documented settings. Consult the linked LiteLLM provider documentation for provider requirements, or follow the README's local-server examples.
- Assemble a small, relevant collection of papers and select its directory. Use the documented
pqaquestion-answering workflow or Pythonaskinterface to build the index and retrieve evidence. - Inspect the formatted answer, citations and evidence context against the source documents. For an intended evaluation, repeat representative questions with controlled settings and record unsupported claims or missed evidence before expanding the collection.