PageIndex and Streamlit: a PDF search app with links to pages
Published: 2026-10-03 · Author: AI Release · @ai_release1
⚡ The gist in 5 seconds - PDF Desk is a Python app for finding answers in a PDF, linking the model's answer to a specific page of the document. - Available on GitHub, runs on macOS/Linux; Windows commands are in the README. - Limitations: one active PDF, up to 25 MiB and 500 pages, a text layer is required, no OCR. ### 🔍 What was found An article on Habr, published October 3, describes PDF Desk — a small app built on Python, Streamlit, PageIndex, and PDFium. The project pins versions: PageIndex 0.2.20, Streamlit 1.64.0, pypdfium2 5.13.0, PyPDF2 3.0.1. It requires Python 3.11 or newer, Git, and access to an OpenAI-compatible POST /chat/completions endpoint. The answer model must support tool calls, otherwise PageIndex cannot complete its workflow. Connection checking is implemented not via GET /models but with two short requests. First, a regular request goes to the preparation model (DeepSeek's deepseek-flash in the test), then the answer model is given a ping function and asked to call it. This catches an incompatible API before the document text is sent. The code is split across three files: app.py (interface and scenario), reader.py (PDFium, PageIndex, validation rules), and storage.py (local paths and the manifest). The index is stored in data/items/<fingerprint>/index/, and the active document in active.json. ### 💡 Why it matters The problem PDF Desk solves is familiar to anyone who has worked with LLM-based chats: the model may confidently say "18 months" where the contract says "12." In a chat, such an error is hard to spot, but here every link in the answer leads to the physical page of the original PDF, which can be opened with one click and verified. It is a practical tool for fact-checking model answers and for building more transparent RAG applications, where an answer always comes with evidence from the primary source. ### 🧩 Context In a previous article, the author covered a local RAG on Markdown and TXT, where the source was a range of lines. A PDF is different: the text may not match the visual representation due to columns, tables, and captions. Therefore, in PDF Desk, verification uses the page as an image generated by PDFium, alongside the text layer to understand what the program saw. PageIndex runs in local mode: it builds a tree index of the PDF and performs agentic search, but model computations may go to an external API — important to keep in mind when working with confidential documents.