Four separate collections, searched together. Each page keeps the locator printed on it, so anything you find can be cited back to the source production.
| NIH-001234 | Fauci e-mail, the Leopold/NIH FOIA release. Bates-stamped.
Pages 680–1603 of that PDF carry no stamp and are cited as p. 812. |
| SLACK_000216 | Slack channel, Bates-stamped. Andersen, Holmes, Rambaut and Garry, April 2020 to June 2023. |
| REC-0042 | Investigative Records, Vol. 1. No stamps in the production, so the locator is the page of that PDF. |
| RR-0198 | Reading Room exhibits, entered at the July 29, 2026 HSGAC hearing. Locator is the page of that PDF. |
Type words and you get pages containing all of them. Everything below is refinement.
furin cleavage site
| AND | Both must appear. This is the default between words, so
bat coronavirus and bat AND coronavirus are the same search. |
| OR | Either. Use it for synonyms and spelling variants, which matter here because much of the text is OCR. |
| NOT | Excludes. -word is a shorthand for the same thing. |
| "…" | An exact phrase. Punctuation and capitals are ignored, so
"gain of function" also finds Gain-of-Function. |
| ( ) | Grouping. (remdesivir OR chloroquine) AND trial. |
| word* | Prefix wildcard. quarantin* catches quarantine, quarantined,
quarantining. |
| NEAR/n | Within n words, in either direction.
lab NEAR/5 leak finds the idea however it is phrased. |
Prefix a word with a field name to search only that part of a document. Fields take one word, so combine them with AND rather than typing a full name.
| from: | Sender. from:daszak |
| to: cc: | Addressee, or copied. The distinction is real and often the point: cc:fauci AND subject:origin |
| subject: | Words in the subject line — the message is about this, rather than merely mentioning it. subject:remdesivir |
| attach: | Words in attachment filenames. attach:gain |
| speaker: | Who is talking in the Slack channel. speaker:andersen AND pangolin |
| corpus: | nih, slack, rec, rr.
"proximal origin" AND corpus:rec |
| year: month: date: | year:2020, month:2020-02,
date:2020-02-01. month:2020-01 AND wuhan |
"gain of function" AND corpus:slack
What the authors of the origins paper said among themselves, as opposed to what was said to them.
furin AND (engineered OR "cleavage site") NOT corpus:rr
The scientific argument, with the hearing exhibits set aside.
from:collins AND (fauci OR nih) AND month:2020-04
One sender, one month. Narrowing by month is usually more effective than adding a third keyword.
wuhan NEAR/10 laboratory
Proximity beats a phrase when you do not know the wording.
hydroxychloroquin* OR chloroquin*
Wildcards absorb OCR damage and word endings at once.
The Fauci e-mail collection is optical character recognition over a scan, and it is imperfect in specific, predictable ways. Words break mid-token, so COVID-19 may sit in the file as COV ID-19. Letters substitute: o for 0, l for 1, NIAID for NIAIO. A search that returns nothing has usually hit one of these rather than proved a negative.
Three habits fix most of it. Search a distinctive fragment rather than a long phrase. Use wildcards on word stems. When a term matters and comes back empty, try it as an OR of two or three plausible manglings before concluding it is absent. The subject-line and sender fields are cleaner than the body text, because they were typed rather than scanned.
The Slack collection and the two Rand Paul volumes have a real digital text layer and behave normally.
A result is a page, not a document. A long thread with an attachment may match on one page and be irrelevant on the next twelve. Open the page and read it.
The text you read here is the extraction, not the document. It cannot show you a redaction:
withheld material appears as (b)(6), as a blank, or as nothing at all. Before you
quote, rely on, or characterise any passage, open the page image or the PDF and look at the
scan. The extraction is a finding aid; the image is the evidence.
Recipient lists in particular are heavily redacted. Absence from a cc: search is
weak evidence that someone was not copied — their name may simply have been withheld.
Every page has its own address. Open a page and press Copy citation — you get the locator, the collection, and a link that opens the archive on that exact page:
fauci-archive.org/#p=NIH-001920
You can build these by hand, and you can link to a search as well as to a page:
| #p=NIH-001920 | opens that page in the reader |
| #q=furin AND corpus:slack | runs that search |
| #q=pangolin&p=NIH-001920 | the search, with the page already open |
| #index #guide #about | a tab |
This is what makes the archive citable rather than merely searchable: a footnote can send a reader to the page itself instead of instructing them to search for it.
Cite the locator, not the page of the PDF: NIH-001920 identifies one page and
one page only, across every copy of this release, forever. Add the archive as the finding aid
and the Internet Archive item as the source of the underlying production. For example:
E-mail, Kristian Andersen to Anthony Fauci, 1 February 2020, NIH-001920, Leopold/NIH FOIA release, Fauci Document Archive, fauci-archive.org.
For the unstamped tranche, where the release carries no page stamp, the locator gives the PDF
page instead (p. 812); say so in the citation, because that number is ours rather
than the producing agency’s.
Search finds strings. The index finds subjects, which is a different thing: it separates a person writing from a person being written about, records the rare single mention that a keyword search buries under a thousand routine ones, and gathers a topic that appears under a dozen different words. Start at the index when you are asking what is in here about X, and at the search when you already know the string you want.
Chip Joyce built this archive — the index, the search tool, and the apparatus around them. On X: @chipjoyce.
The permanent address is fauci-archive.org. Cite that, not any hosting provider’s address: the archive can move hosts without breaking a single citation, which is the point of owning the domain.
It exists because a 6,228-page pile of released documents is not the same thing as an available record. The documents were public before this site; they were not usefully findable. The work here is the finding aid: an index built by editorial judgment rather than machine frequency, a Boolean search over the full text, and a locator scheme that lets anything found here be cited back to the page it came from.
Four separate public releases, reproduced as released, with nothing added, removed, or reordered:
| NIH- | Anthony Fauci’s e-mail, obtained under the Freedom of Information Act and published in 2021 following requests by Jason Leopold and others. |
| SLACK_ | A Slack channel among Kristian Andersen, Edward Holmes, Andrew Rambaut and Robert Garry, released by Chairman Rand Paul. |
| REC- | Investigative Records, Vol. 1, released by Chairman Rand Paul. |
| RR- | Reading Room exhibits entered into the record at the 29 July 2026 hearing of the Senate Committee on Homeland Security and Governmental Affairs. |
This is an independent project. It is not affiliated with, endorsed by, or produced in cooperation with the National Institutes of Health, NIAID, the Department of Health and Human Services, any Senate or House committee, any member of Congress, or any person named in these documents.
Inclusion here carries no implication about anyone. An index heading is a finding aid, not an allegation; a document appearing under a topic means the words are on the page, not that any claim about the topic is true. Readers draw their own conclusions from the documents.
The text is imperfect and in places wrong. Most of the Fauci e-mail collection is optical character recognition over a scan, which breaks words, substitutes letters, and occasionally loses passages outright. Documents in these productions appear out of sequence, some pages carry no page stamp, and material has been withheld under statutory exemptions. Extracted text cannot show you a redaction. Before quoting, relying on, or characterising any passage, check it against the page image or the source PDF. The extraction is a finding aid; the scan is the evidence.
The editorial judgments — which subjects earn a heading, which mentions earn a locator,
which are set aside as too frequent to list — are judgments, and reasonable people would
make some of them differently. They are written down in STYLE_SHEET.md so they can
be inspected and disputed rather than merely trusted.
Nothing here is legal advice, and the summary of rights below is a description of how this work is offered, not a legal opinion about the underlying documents.
Two different things are stacked here, and they do not share a copyright status.
The documents. These are public records obtained through FOIA and through congressional release. Works of the United States government are not subject to copyright in the US. Some material within these productions was written by private individuals and organisations — the Slack channel most obviously — and its status is whatever it is; that status is not changed by its appearing in a public release, and it is not affected one way or the other by anything on this site. No claim of copyright is asserted over any of these documents here. They are reproduced because the public has a right to read what was released.
The apparatus. The index, the controlled vocabulary and style sheet, the document registers, the locator scheme, the search engine and the site are original work by Chip Joyce. They are dedicated to the public domain under CC0 1.0. Copy them, host them, fork them, build on them, sell something built on them, with or without credit and with or without asking. A finding aid that researchers cannot freely mirror is a finding aid with a single point of failure.
Attribution is not required. It is appreciated, and it helps readers trace a citation back to the machinery that produced it.
All four productions are archived at the
Internet Archive,
together with a complete copy of this site. Three are downloadable individually; the Slack
production is preserved inside fauci-pdfs-originals.zip, which also holds all
four files exactly as released, byte for byte, before any structural repair.
Every page of every production is reproduced here as an image, so a reader rarely needs the PDFs. The scan shows redactions; the extracted text cannot.
The whole archive is static files. There is no server, no database, no tracking, and no analytics of any kind. Copy the folder onto any host and it works; open it from a hard drive and it works. The parsed data sits beside the site as plain JSON — page text, page metadata, the document register, the message register, the index with locators resolved — so anyone who wants to build a different tool over the same corpus can start from those files rather than re-deriving them from the PDFs.
Corrections are welcome, particularly misattributed documents, mangled names, and index headings that point somewhere they should not.