Corrupt & Invalid File Generator
Download test files that are broken in exactly one, documented way — 130 cases across 15 formats, each checked against real parsers — to see how your upload validation and file parsers fail. Everything is generated in your browser.
Truncated (cut in half)Some readers reject
A valid file cut off halfway, like an interrupted upload or download.
Tests: Detection of incomplete transfers and missing end markers.
Some readers reject it; more forgiving ones accept or repair it.
Nothing generated yet. Pick a format and a defect, then download it — or grab the whole test pack for the format.
One-click broken files
Click to download; the generator above fills in too, so you can change the size or seed and download again.
Why Test With Broken Files?
Most file-handling bugs live in the error path. An upload form that accepts a valid PDF tells you little; what matters is what happens when someone sends half a PDF, a PNG renamed to .pdf, a password-protected document or a server's 404 page saved as invoice.xlsx. Each file here is broken in one documented way, so when a test fails you know exactly which check is missing.
- Upload validation — does your form check the content (magic bytes) or only trust the extension and the browser's MIME type?
- Parsers and converters — do PDF, image and Office libraries raise a clean, catchable error, or crash the worker?
- Error messages — does the user see "this file is damaged" or a 500 page?
- Security scanners and queues — do password-protected and malformed files get quarantined, retried forever, or silently dropped?
What the Result Types Mean
Every case was opened with real libraries (listed below). The result type says what they did, so you know what to expect from yours:
- Unreadable — every reader we tested rejects it.
- Some readers reject — strict parsers fail while forgiving ones repair or accept it. A truncated PDF is the classic example:
pypdfin strict mode refuses it, MuPDF and browsers rebuild what they can. - Opens anyway — most software opens it, so only a validator that checks for this exact problem catches it.
- Password — a valid file that can't be opened without the password (it's
testfor every protected sample). - Valid edge case — a correct file with something your code may not expect: a PDF with no pages, an empty ZIP, a PDF that forbids copying.
Every Defect, Format by Format
15 formats, 45 kinds of defect, 130 combinations. Click any defect to load it into the generator.
PDF 12 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Some readers reject |
| PNG renamed .pdf | A real file of another type saved with this extension. | Unreadable |
| HTML error page | A "404 Not Found" web page saved with this extension, the classic broken download. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| Missing %%EOF marker | The end-of-file marker is gone, so the file looks incomplete. | Some readers reject |
| Broken cross-reference table | Every xref entry points at the wrong object. Readers rebuild the index silently (MuPDF reports it as repaired). | Opens anyway |
| Wrong stream /Length | The page’s content stream declares the wrong length. | Some readers reject |
| No pages | A structurally valid PDF whose page tree is empty. | Valid edge case |
| Password-protected (password: test) | RC4 128-bit encryption with the open password "test". | Password |
| Restricted (no print or copy) | Opens without a password, but an owner password forbids printing, copying and editing. | Valid edge case |
Word (DOCX) 9 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Unreadable |
| PNG renamed .docx | A real file of another type saved with this extension. | Unreadable |
| HTML error page | A "404 Not Found" web page saved with this extension, the classic broken download. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| CRC mismatch in a part | The main document part fails its ZIP checksum. | Unreadable |
| Missing [Content_Types].xml | The package has no content-types part, so it isn’t a valid Office file. | Unreadable |
| Malformed XML inside | The ZIP is intact, but the main XML part has a mismatched closing tag. | Unreadable |
Excel (XLSX) 9 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Unreadable |
| PNG renamed .xlsx | A real file of another type saved with this extension. | Unreadable |
| HTML error page | A "404 Not Found" web page saved with this extension, the classic broken download. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| CRC mismatch in a part | The main document part fails its ZIP checksum. | Unreadable |
| Missing [Content_Types].xml | The package has no content-types part, so it isn’t a valid Office file. | Unreadable |
| Malformed XML inside | The ZIP is intact, but the main XML part has a mismatched closing tag. | Unreadable |
PowerPoint (PPTX) 9 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Unreadable |
| PNG renamed .pptx | A real file of another type saved with this extension. | Unreadable |
| HTML error page | A "404 Not Found" web page saved with this extension, the classic broken download. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| CRC mismatch in a part | The main document part fails its ZIP checksum. | Unreadable |
| Missing [Content_Types].xml | The package has no content-types part, so it isn’t a valid Office file. | Unreadable |
| Malformed XML inside | The ZIP is intact, but the main XML part has a mismatched closing tag. | Unreadable |
ZIP archive 9 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Unreadable |
| PNG renamed .zip | A real file of another type saved with this extension. | Unreadable |
| HTML error page | A "404 Not Found" web page saved with this extension, the classic broken download. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| CRC mismatch | An entry’s stored CRC-32 doesn’t match its data. | Unreadable |
| Password-protected (password: test) | Traditional ZIP encryption (ZipCrypto) with the password "test". | Password |
| Empty archive (no files) | A valid ZIP with no files in it (padded with an archive comment). | Valid edge case |
PNG image 9 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Some readers reject |
| PDF renamed .png | A real file of another type saved with this extension. | Unreadable |
| HTML error page | A "404 Not Found" web page saved with this extension, the classic broken download. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| Bad chunk checksum | The IHDR chunk’s CRC-32 does not match its data. | Unreadable |
| Corrupt pixel data | The compressed image data is damaged; the chunk checksums are fixed up so only decoding fails. | Some readers reject |
| Zero width | The header declares a 0-pixel-wide image (with a valid checksum). | Unreadable |
JPG image 8 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Some readers reject |
| PDF renamed .jpg | A real file of another type saved with this extension. | Unreadable |
| HTML error page | A "404 Not Found" web page saved with this extension, the classic broken download. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| Missing end marker | The final FF D9 end-of-image marker is missing. | Some readers reject |
| Zero height | The frame header declares a 0-pixel-high image. | Unreadable |
GIF image 8 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Some readers reject |
| PDF renamed .gif | A real file of another type saved with this extension. | Unreadable |
| HTML error page | A "404 Not Found" web page saved with this extension, the classic broken download. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| Unknown GIF version | The signature reads GIF89x instead of GIF89a. | Unreadable |
| Missing trailer | The closing 0x3B trailer byte is missing. | Opens anyway |
TIFF image 7 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Some readers reject |
| PDF renamed .tif | A real file of another type saved with this extension. | Unreadable |
| HTML error page | A "404 Not Found" web page saved with this extension, the classic broken download. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| IFD offset past the end | The header points to an image directory beyond the end of the file. | Unreadable |
WAV audio 8 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Opens anyway |
| PNG renamed .wav | A real file of another type saved with this extension. | Unreadable |
| HTML error page | A "404 Not Found" web page saved with this extension, the classic broken download. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| Unknown audio format code | The fmt chunk declares codec 0x9999 instead of PCM. | Unreadable |
| Data size larger than file | The data chunk claims twice the bytes actually present. | Opens anyway |
JSON 13 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Unreadable |
| PNG renamed .json | A real file of another type saved with this extension. | Unreadable |
| HTML error page | A "404 Not Found" web page saved with this extension, the classic broken download. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| Trailing comma | A comma before the closing brace: {"a":1,}. | Unreadable |
| Single-quoted key | A key written as 'meta' instead of "meta". | Unreadable |
| Missing closing brace | The top-level object is never closed. | Unreadable |
| Comment inside | A /* comment */, which JSON doesn’t allow. | Unreadable |
| Byte-order mark | Starts with a UTF-8 BOM (EF BB BF). | Some readers reject |
| NaN value | A bare NaN, which JSON doesn’t allow. | Some readers reject |
| Invalid UTF-8 | Invalid UTF-8 bytes inside a string. | Some readers reject |
XML 9 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Unreadable |
| PNG renamed .xml | A real file of another type saved with this extension. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| Unclosed root element | The closing root tag is missing. | Unreadable |
| Undefined entity | Uses &bogus;, an entity that was never declared. | Unreadable |
| Two root elements | A second element after the root closes. | Unreadable |
| Illegal control character | Contains a 0x01 byte, which XML 1.0 forbids everywhere. | Unreadable |
SVG 8 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Unreadable |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Unreadable |
| PNG renamed .svg | A real file of another type saved with this extension. | Unreadable |
| Right header, random body | Starts with the correct signature bytes, then random data. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| Unclosed root element | The closing root tag is missing. | Unreadable |
| Illegal control character | Contains a 0x01 byte, which XML 1.0 forbids everywhere. | Unreadable |
| Missing SVG namespace | Well-formed XML, but without xmlns, so browsers don’t treat it as SVG. | Opens anyway |
CSV 7 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Some readers reject |
| Truncated (cut in half) | A valid file cut off halfway, like an interrupted upload or download. | Opens anyway |
| PNG renamed .csv | A real file of another type saved with this extension. | Unreadable |
| Random bytes | Pure random data with no structure at all. | Unreadable |
| Row with an extra column | One row has 7 fields where the header has 6. | Some readers reject |
| Unclosed quote | The last row opens a quoted field that never closes. | Some readers reject |
| Invalid UTF-8 | Invalid UTF-8 bytes in a row. | Unreadable |
TXT (UTF-8 text) 5 defects
| Defect | What is broken | Result |
|---|---|---|
| Empty file (0 bytes) | A zero-byte file with the right extension. | Valid edge case |
| PNG renamed .txt | A real file of another type saved with this extension. | Some readers reject |
| Random bytes | Pure random data with no structure at all. | Some readers reject |
| Invalid UTF-8 | Invalid UTF-8 byte sequences scattered through the text. | Some readers reject |
| NUL bytes | Valid UTF-8 with 0x00 bytes inside, which most tools treat as binary. | Opens anyway |
How Each File Is Checked
The result types aren't guesses. Every combination was generated at five or more sizes, from 1 byte to 600 KB, and opened with these readers on 10 October 2026: pypdf 6.14 in strict mode and PyMuPDF 1.28 for PDF; Pillow 12.1 (strict, then with truncated images allowed) for PNG, JPG, GIF and TIFF; Python 3.11's zipfile, wave, json and xml.etree; lxml 6.1; pandas 3.0 and Python's csv module; python-docx 1.2, openpyxl 3.1 and python-pptx 1.0 for Office files; and JSON.parse in Node 22 and Microsoft Edge. Your libraries may be stricter or more forgiving, which is exactly what these files help you find out.
Using a Test Pack in CI
Download test pack builds one ZIP per format: a valid control file, one file per defect, a README.txt and a manifest.json that lists each file's defect, result type, size, SHA-256 and password. Commit it as fixtures and drive a parametrised test from the manifest:
import json, pathlib, pytest
from myapp.uploads import validate_upload, InvalidFile
pack = pathlib.Path("fixtures/pdf")
cases = json.loads((pack / "manifest.json").read_text())
@pytest.mark.parametrize("case", cases, ids=lambda c: c["defect"])
def test_upload(case):
data = (pack / case["file"]).read_bytes()
if case["kind"] == "valid":
validate_upload(data)
elif case["kind"] in ("unreadable", "strict"):
with pytest.raises(InvalidFile):
validate_upload(data)
Add a seed and the pack is reproducible byte for byte, so you can regenerate it in another branch and compare hashes.
What This Tool Deliberately Doesn't Make
- No malware or antivirus test files. If you need to check that a scanner fires, use the industry-standard test file from EICAR instead.
- No zip bombs, XML entity bombs or XXE payloads, path-traversal archives, macros or executables. Those are attack payloads, not QA fixtures, and they belong in a controlled security test.
- No "corrupt my file" upload. Every file is synthetic. The tool won't damage a document you already have — that is mostly used to send someone a broken file on purpose.
Need files that are valid at an exact size instead? Use the Sample File Generator.
Frequently Asked Questions
test. The PDF uses the standard RC4 128-bit security handler and the ZIP uses traditional PKWARE encryption (ZipCrypto), the two kinds most upload pipelines meet. The "restricted" PDF opens without a password but sets an owner password that forbids printing, copying and editing.python-docx, openpyxl and python-pptx raise an error for every DOCX, XLSX and PPTX case here: a bad checksum fails the ZIP integrity check, a missing [Content_Types].xml means the package isn't valid Office Open XML, and the malformed-XML case fails when the main part is parsed. Desktop Office may offer to repair some of them, which is worth testing too if your users open files in Office.csv module reads a row with an extra column. Those cases are marked Some readers reject or Opens anyway; they test whether your validation is stricter than the reader you happen to use.