Initializing secure environment…
Initializing secure environment…
Batch documents arrive as one long file, and the only reliable boundary is the text itself. Give this tool a phrase — an invoice number, a heading, a date line — and it starts a new file at every match, in order. Matches are shown with a page and a text preview before you commit, so a heading repeated in a footer can be spotted and excluded. Nothing is uploaded, there is no watermark and no sign-up. It needs a text layer, so OCR a scan first.
To split a PDF by text, add the file and type the phrase that starts each new document, such as an invoice number or a heading. Every match starts a new PDF, listed with a preview so you can confirm the boundaries before downloading.
Use a phrase unique to each document — the invoice number, or 'Tax Invoice' if your template is consistent. Then exclude the match that appears in the running total page. The result is one PDF per invoice, named from the matched line, ready to be zipped and emailed.
Statements repeat the account holder's name on every page, so a heading phrase is better than a name. If the statement has a month heading on its first page, that phrase gives exactly one boundary per statement.
A contents page lists every chapter heading, so every heading appears once on the contents page and once at the chapter. Exclude the contents page from the match list, or split starting after it, and the boundaries are clean.
When several unrelated PDFs were merged by mistake, the first line of each original is a reliable boundary. Split by that phrase and the pack separates into exactly the files that went in.
Anything generated in batches that repeat a fixed phrase: invoice runs, bank statements, dunning letters, signed engagement letters, exam question papers. The phrase has to appear in a text layer, so a document produced by a system that outputs real text works and a pure scan does not.
Not until the scan has a text layer. Run OCR first, then split by text on the result. OCR is the honest prerequisite — matching against pixels would either fail silently or produce boundaries nobody could defend.
As text, not as a pattern. Case and surrounding whitespace are normalised, so INVOICE and Invoice both match. Regular expressions and wildcards are not supported, because a document boundary that depends on a regular expression is a boundary nobody can audit.
They become their own file, named for the preamble, so a cover page or a summary sheet is not silently discarded. If you would rather drop them, exclude that first file before downloading.
A boundary cannot be finer than a page, so the match that appears later in the page ordering starts the next file. The preview shows the page text so you can see when a heading is not at the top of its page.
Almost always a phrase that also appears in a footer, a page header, or a summary line. The match list shows page and preview; exclude those and split again. Narrowing the phrase to something unique to each document is the other fix.
Yes. Pages are copied rather than re-rendered, so fonts, the text layer, images and links are all preserved. The files print exactly as the original pages did.
Free, no sign-up, no watermark. The document is only read inside this tab, so a batch of tax documents never reaches a server.
More essential pdf tools — all free, no upload.