Parquet, Avro & Arrow File Viewer
See what is inside a Parquet, Avro or Arrow file — no Python, nothing uploaded.
Open a Parquet, Avro or Arrow file
Contents
Columns
Value
Columns in file order, with nested fields indented. The type in bold is how this viewer reads the column; the grey text is the type as the file declares it.
The schema as the file writes it
Row groups
Columns
Min and max are the statistics the writer stored. “≥” and “≤” mark shortened values (bounds, not the exact minimum or maximum); an empty cell means none were stored.
Metadata stored in the file
This file has no key/value metadata.
About the Parquet, Avro & Arrow File Viewer
Open an Apache Parquet, Avro or Arrow file — the formats data pipelines, Spark, pandas, DuckDB, BigQuery exports and machine-learning datasets use — and look inside without installing Python or a database. Browse the rows page by page, choose which columns to show, and open any value in full: nested lists, maps and structs are shown as JSON, decimals exactly, timestamps to the micro- or nanosecond.
The Schema tab shows every column with its type, including nested fields, and the schema as the file writes it. File details shows what a file is made of: rows, row groups (Parquet), blocks (Avro) or record batches (Arrow), compression codecs, sizes, and for Parquet the statistics of every column chunk — min, max, null count — and its encodings. When you have what you need, export the rows to CSV, Excel, JSON or JSON Lines. The file is read in pieces straight from your disk, so even multi-gigabyte files open quickly; nothing is uploaded.
How to use it
- Drop a
.parquet,.avro,.arrow,.featheror.arrowsfile on the box, or choose it. Or press Open a sample Parquet file to try the viewer on 3,000 made-up orders. - Data shows the first rows. Move through the pages at the bottom, or change Rows per page. Use Columns to show only the columns you need (with a search box for wide files).
- Select a cell and press Enter (or double-click it) to see the whole value — long text, a list, a map, a struct — with a Copy button; Esc takes you back to the cell.
- Open Schema for the column types and nesting, and File details for the row groups, compression, statistics and the metadata stored in the file. For a Parquet row group, press Columns to see each column chunk.
- To export, pick a format (CSV, Excel, JSON, JSON Lines or TSV) and All rows or Rows on this page, then press Download. Only the columns shown are exported.
Examples
Open a sample Parquet file → Data, row 1
order_id 100001 · customer Chen Wei · country GB · items ["Plant pot", "Desk lamp", "Notebook"] · quantity 5 · total 51.48 · paid true · ship_to {"city": "Leeds", "postcode": "LS1 4AP"} · note NULLThe sample has 3,000 rows in three row groups of 1,000, Snappy-compressed. File details shows each row group’s size and, under Columns, every column chunk’s min, max and null count.
sales.parquet, 250,000 rows · Export: Excel (.xlsx) · All rows
sales.xlsx with a bold, frozen header row; numbers stay numbers and timestamps become Excel dates
Integers and decimals with more than 15 significant digits are written as text, because Excel keeps only 15 digits of a number.
Common uses
- Checking what a Parquet export from Spark, pandas, DuckDB, BigQuery or a data lake contains before loading it anywhere.
- Looking at the schema of an Avro file from Kafka, Hadoop or a data pipeline, nested records and unions included.
- Opening a Feather or Arrow IPC file shared by a colleague who works in Python or R.
- Turning a Parquet or Avro file into CSV or Excel for someone who does not use data tools.
- Seeing why a Parquet file is large: row groups, codecs, encodings and compression ratios per column.
What the viewer reads
- Parquet: data pages version 1 and 2; the codecs UNCOMPRESSED, SNAPPY, GZIP (also several gzip members in one page), ZSTD, BROTLI, LZ4_RAW and the deprecated LZ4 (Hadoop framing or a plain block), as listed in Parquet’s compression codecs; the encodings PLAIN, dictionary (RLE_DICTIONARY and PLAIN_DICTIONARY), RLE, DELTA_BINARY_PACKED, DELTA_LENGTH_BYTE_ARRAY, DELTA_BYTE_ARRAY and BYTE_STREAM_SPLIT (Parquet encodings); nested lists, maps and structs; the logical types for strings, decimals, dates, times, timestamps (milli-, micro- and nanoseconds), unsigned and small integers, UUID, JSON, FLOAT16 and the old INT96 timestamps. Parquet is read with hyparquet (MIT); the Zstandard (RFC 8878), Brotli (RFC 7932) and LZ4 decoders are this site’s own.
- Avro object container files: the codecs null, deflate, snappy and zstandard; records, enums, arrays, maps, unions, fixed, recursive types and namespaces, and the logical types decimal, uuid, date, time, timestamp, local-timestamp and duration from the Avro specification.
- Arrow IPC files (also called Feather version 2,
.arrowor.feather) and IPC streams (.arrows), with buffers compressed by LZ4 or Zstandard, dictionaries (including delta dictionaries), unions, run-end encoded columns, string and binary views, and list views (Arrow columnar format).
How values are shown and exported
Decimals keep every digit (12.34, not 12.3399999). Timestamps are written in ISO 8601 with the precision of the column — YYYY-MM-DDTHH:MM:SS.ffffffZ for microseconds — ending in Z (UTC) when the file marks the column as adjusted to UTC, and without a zone when it stores local (“wall clock”) times. Dates are YYYY-MM-DD; times of day HH:MM:SS with their fraction; durations carry their unit (90 s). Integers are exact at any size. Binary values appear as hexadecimal (0x0a1b…) in the grid, CSV and Excel, and as Base64 in JSON. Lists, maps and structs become JSON.
In JSON exports, integers beyond 2^53 and decimals are written as strings, so programs that read JSON numbers as floating point do not lose digits. Excel holds 1,048,576 rows per sheet and 15 significant digits per number (Microsoft: Excel specifications and limits): longer numbers are written as text, and an Excel export stops after 1,048,575 rows (the export says so — use CSV or JSON Lines for more). Excel’s calendar starts in 1900 and counts 1900 as a leap year (Microsoft Learn), so dates in the first two months of 1900, and all earlier dates, are written as text rather than as Excel dates.
Big files
A Parquet file keeps its metadata at the end (Parquet file format): the viewer reads that footer first, then only the column chunks of the row group on screen, for the columns shown — the rest of the file is never read. Avro blocks are indexed on open without being decompressed, and Arrow record batches are read one at a time. So a page of rows appears quickly even from a file of several gigabytes. What has to fit in memory is one row group or batch of the columns you show: if a file has very large row groups, show fewer columns.
Limitations
- Parquet columns compressed with LZO, stored with the new ALP encoding, or stored with DELTA_BYTE_ARRAY in version 1 data pages are shown as unreadable — the other columns, the schema and the file details still open. Files encrypted with Parquet modular encryption cannot be opened.
- Avro files compressed with bzip2 or xz show their schema and metadata but not their rows.
- Feather version 1 files (the format before Arrow IPC), big-endian Arrow files and Apache ORC files are not supported.
- There is no sorting or filtering of rows: export to CSV and use Run SQL on CSV & Excel files for that.
Privacy
Everything happens in your browser. What you enter or open here is not uploaded or stored by MySmartCoPilot. Your file is read on your device and never uploaded. Only for Parquet files compressed with Brotli is the Brotli dictionary (about 80 KB) downloaded from this site, once.
Frequently asked questions
How do I open a Parquet file without Python?
Drop it on this page. The schema, the row count and the first rows appear in a moment, and you can page through the rest, look at any value in full, or export everything to CSV or Excel. Nothing to install, and the file stays on your device.
Is my file uploaded?
No. The file is read by code running in your browser, in pieces, straight from your disk. Nothing is sent anywhere — the page also works offline once it has loaded.
How do I convert Parquet to CSV or Excel?
Open the file, choose the columns you want under Columns (all are shown at first), pick CSV or Excel (.xlsx) next to Export, and press Download. The whole file is converted, a row group at a time, so large files work too. Avro and Arrow files export the same way.
What is a row group?
Parquet stores a table in horizontal slices called row groups; inside each, every column is stored separately as a column chunk, compressed and with its own statistics (min, max, null count). Readers use the statistics to skip row groups they do not need. File details lists the row groups with their sizes, and Columns shows each column chunk’s codec, encodings and statistics.
Can it open Feather files?
Yes — Feather version 2 is the Arrow IPC file format, which this viewer reads (with or without LZ4 or Zstandard compression). The older Feather version 1 is not supported: open it in Python and save it again with pyarrow.feather.write_feather, which writes version 2.
Why do some timestamps end in Z and others not?
Parquet, Avro and Arrow say whether a timestamp is an instant in UTC or a local date and time with no time zone. UTC instants end in Z; local ones are shown as stored, without a zone, because the file does not say which zone they belong to.