Skip to main content
Start with a file ID from search or the file inventory. Use connection setup for authentication and request conventions.
The response contains files and nextCursor. Select a returned files[].fileId for the content requests below. Follow inventory cursors until null if you need every caller-visible stored file. Select a document using its role, name, and research relevance, then verify its contents. Stored filenames and search passages alone do not establish the full meaning of a form or policy.

Read existing extracted text

An abbreviated response fragment:
A success returns text, textRevision, byte location, and nextCursor, with the file ID and a nullable source versionId. See the text reference for the schema and size bounds. Byte offsets count UTF-8 bytes, not characters; a chunk can be shorter than maxBytes to preserve a complete character. This reads existing extraction only. A downloadable file or indexed search match need not have readable stored text. 409 text_unavailable permits a fallback to the original download; an available empty extraction instead succeeds with empty text and a null cursor. 404 not_found is not the same as missing extraction. See recovery for access and environment failures.

Save a checkpoint and resume

Save the file ID, maxBytes, textRevision, and the cursor used to request a chunk alongside its text, byte offsets, and returned nextCursor. Store the chunk before advancing the checkpoint. When nextCursor is non-null, resume using it:
Keep the same file and size. If a response is lost, retry the same input request; deduplicate saved chunks by file ID, revision, and starting byte offset. Returned cursor strings may differ on retry. Stop when nextCursor is null. Cursors expire 24 hours after the first chunk; continuation does not extend their lifetime, and every request checks access again. On invalid_cursor, restart the intended text traversal separately. Never append a new revision to old chunks. See pagination and recovery for failure handling.

Download the original

Use the download operation when you need layout, original pages, or a file without readable extraction. Use the inventory’s versionId when present. If it is null, omit that parameter to pin the current original version.
The redirect identifies the selected version:
Follow the signed URL without forwarding API authorization to the storage host. After the download completes, inspect the actual bytes. Keep File-Version-Id for citations and retries; the signed URL expires and is not a durable reference. To retry, request the same file and saved version ID for a fresh redirect. Do not silently substitute current bytes when an exact version is unavailable. Use bounded concurrency if downloading multiple files.

Cite what you actually read

The text operation currently returns source versionId: null; its revision identifies extracted text only. Embedded page markers do not establish verified page mapping. For a PDF citation, inspect the actual downloaded page and record that original’s version. A local PDF-to-text conversion derives from those local bytes, but it is not an API textRevision. For example: “Read page 1 of the downloaded form, file ID X, version V; search evidence version was unknown.” That supports a source-grounded statement without inventing a version relationship. Keep extracted research data and your interpretation separate from the original evidence.