Back to all posts
Tutorial
8 min read

How to Read and Open a Parquet File

DevToolLab Team

DevToolLab Team

October 11, 2026

How to Read and Open a Parquet File

To read a Parquet file in Python, install pyarrow and call pd.read_parquet("orders_2025.parquet"). From a terminal, duckdb -c "SELECT * FROM 'orders_2025.parquet' LIMIT 3" prints the first rows with their column types. To open one without code on Windows or a Mac, drop it into a browser Parquet viewer. A text editor such as Notepad++ only shows binary noise.

The thing to know is that Parquet is a binary, column-oriented format. According to the Apache Parquet file format spec, a file starts and ends with the 4-byte marker PAR1, and the schema and row group locations sit in a footer just before the last marker. Every reader seeks to the end first, then pulls only the columns you ask for. That is why reading two columns is cheap, and why a half-downloaded file fails outright.

This guide follows one file the whole way: orders_2025.parquet, a 1,200,000-row, 7-column US order export shaped like the one in our large CSV guide, written by DuckDB 1.5.6. It weighs 15.6 MB, against 67.8 MB for the same data as CSV, and holds 10 row groups of 122,880 rows compressed with Snappy. Every output below was produced on October 11, 2026, on an Apple M3 Pro with 18 GB of RAM.

The Fastest Way

Drop the file into the Parquet Viewer. It reads the footer first and decodes column chunks in the browser tab, so orders_2025.parquet showed its schema, its 10 row groups and the first 50 of 1,200,000 rows after 4,057 ms in Chrome, with nothing uploaded. From there you can search every row and export all rows, or only the matches, as CSV or JSON. It decodes Snappy, Gzip, ZSTD and LZ4 pages; a file compressed with Brotli or LZO has to be re-saved first.

Parquet Viewer with orders_2025.parquet open: 1,200,000 rows, 7 columns, 10 row groups and SNAPPY compression, above the first rows of order_id, order_date, customer_name, city, state, zip and amount_usd with each column's type
Parquet Viewer with orders_2025.parquet open: 1,200,000 rows, 7 columns, 10 row groups and SNAPPY compression, above the first rows of order_id, order_date, customer_name, city, state, zip and amount_usd with each column's type

When the file is part of a script or a pipeline, pick a reader by where the next step happens:

ReaderInstallRead callMeasured on orders_2025.parquet
pandas 3.0.6pip install pandas pyarrowpd.read_parquet(path)full read 0.02 to 0.09 s, 2 columns 0.005 s
pyarrow 26.0.0pip install pyarrowpq.read_table(path)Texas filter 0.006 to 0.009 s
Polars 2.0.0pip install polarspl.scan_parquet(path)December query 0.004 s warm
DuckDB 1.5.6one binarySELECT * FROM 'path'state totals 0.01 to 0.02 s, including startup
PySpark 4.2.0pip install pyspark plus a JDKspark.read.parquet(path)5.5 s, mostly JVM startup

The Python timings are the read call alone, best and worst of five runs; the DuckDB and Spark ones include starting the process.

How to Read a Parquet File in Python With pandas

pandas has no Parquet decoder of its own. It hands the file to pyarrow, or to fastparquet if pyarrow is missing, so install pyarrow alongside it:

Python
import pandas as pd

df = pd.read_parquet("orders_2025.parquet")
print(df.shape)
print(df.dtypes)
print(df[df.city == "Boston"].head(2))
text
(1200000, 7)
order_id           int64
order_date        object
customer_name        str
city                 str
state                str
zip                  str
amount_usd       float64
dtype: object
   order_id  order_date   customer_name    city state    zip  amount_usd
1    100001  2025-05-01  Brandon Taylor  Boston    MA  02108      381.90
5    100005  2025-03-12    Rachel Lopez  Boston    MA  02108        5.67

Two details show up right away. The Boston ZIP code stays 02108, because Parquet stored the column as a string. Running pd.read_csv on the CSV copy of the same data turned that column into integers and printed 2108. And order_date arrives as object, a column of Python date values; pass dtype_backend="pyarrow" to keep it as Arrow's date32[day] instead.

The full read took 0.02 to 0.09 seconds and 139 MB of memory, while pd.read_csv on the CSV took 0.43 to 0.48 seconds. When you only need a few columns, name them, and the rest are never decoded:

Python
import pandas as pd

df = pd.read_parquet("orders_2025.parquet", columns=["state", "amount_usd"])
print(df.groupby("state")["amount_usd"].sum().nlargest(3).round(2))
text
state
TX    60458222.43
CA    30188386.67
OR    15314491.16
Name: amount_usd, dtype: float64

That read took 0.005 seconds.

Read Parquet Without pandas Using pyarrow

To check the schema and row group layout before loading anything, or to filter rows while the file is scanned, call pyarrow directly:

Python
import pyarrow.parquet as pq

f = pq.ParquetFile("orders_2025.parquet")
print(f.metadata)
print(f.schema_arrow)

tx = pq.read_table("orders_2025.parquet", columns=["order_id", "amount_usd"], filters=[("state", "=", "TX")])
print(tx.num_rows, tx.column("amount_usd").to_pylist()[:3])
text
<pyarrow._parquet.FileMetaData object at 0x1035c3ab0>
  created_by: DuckDB version v1.5.6 (build 069cc9f9b5)
  num_columns: 7
  num_rows: 1200000
  num_row_groups: 10
  format_version: 1.0
  serialized_size: 6804
order_id: int64
order_date: date32[day]
customer_name: string
city: string
state: string
zip: string
amount_usd: double
239088 [282.96, 25.14, 411.7]

pyarrow reads that metadata from the 6,804-byte footer without touching the row data. The filters argument drops rows that do not match during the scan, so the full table never sits in memory. Here that returned 239,088 Texas orders.

Query It Lazily With Polars

Polars' scan_parquet builds a query plan first and reads only what the final result needs:

Python
import polars as pl

pl.Config.set_fmt_float("full")

top = (
    pl.scan_parquet("orders_2025.parquet")
    .filter(pl.col("order_date") >= pl.date(2025, 12, 1))
    .group_by("state")
    .agg(pl.len().alias("orders"), pl.col("amount_usd").sum().round(2).alias("revenue_usd"))
    .sort("revenue_usd", descending=True)
    .head(3)
    .collect()
)
print(top)
text
shape: (3, 3)
┌───────┬────────┬─────────────┐
│ state ┆ orders ┆ revenue_usd │
│ ---   ┆ ---    ┆ ---         │
│ str   ┆ u32    ┆ f64         │
╞═══════╪════════╪═════════════╡
│ TX    ┆ 20545  ┆ 5212198.1   │
│ CA    ┆ 10193  ┆ 2588785.11  │
│ CO    ┆ 5211   ┆ 1319827.67  │
└───────┴────────┴─────────────┘

The December query ran in 0.004 seconds once warm and 0.11 seconds on the first call. Nothing happens until collect(), which is what lets Polars skip the four columns the query never mentions.

How to Read a Parquet File From the Terminal With DuckDB

DuckDB 1.5.6, released September 28, 2026, ships its CLI as a single binary and treats a Parquet path as a table:

Bash
duckdb -c "SELECT * FROM 'orders_2025.parquet' LIMIT 3"
text
┌──────────┬────────────┬────────────────┬──────────┬─────────┬─────────┬────────────┐
│ order_id │ order_date │ customer_name  │   city   │  state  │   zip   │ amount_usd │
│  int64   │    date    │    varchar     │ varchar  │ varchar │ varchar │   double   │
├──────────┼────────────┼────────────────┼──────────┼─────────┼─────────┼────────────┤
│   100000 │ 2025-06-11 │ Emily Smith    │ New York │ NY      │ 10001   │     368.78 │
│   100001 │ 2025-05-01 │ Brandon Taylor │ Boston   │ MA      │ 02108   │      381.9 │
│   100002 │ 2025-06-27 │ Michael Davis  │ Miami    │ FL      │ 33130   │     254.59 │
└──────────┴────────────┴────────────────┴──────────┴─────────┴─────────┴────────────┘

The second header row is the schema, so this one command answers both "what is in this file" and "what types are the columns". An aggregate over all 1.2 million rows took 0.01 to 0.02 seconds across three runs, startup included. DuckDB's Parquet docs call this projection pushdown: only the columns a query needs are read from the file.

Bash
duckdb -c "SELECT state, count(*) AS orders, round(sum(amount_usd), 2) AS revenue_usd
FROM 'orders_2025.parquet' GROUP BY state ORDER BY revenue_usd DESC LIMIT 3"
text
┌─────────┬────────┬─────────────┐
│  state  │ orders │ revenue_usd │
│ varchar │ int64  │   double    │
├─────────┼────────┼─────────────┤
│ TX      │ 239088 │ 60458222.43 │
│ CA      │ 119777 │ 30188386.67 │
│ OR      │  60387 │ 15314491.16 │
└─────────┴────────┴─────────────┘

How to Open a Parquet File in Excel

Excel does not list Parquet among its documented data sources. Microsoft's Parquet connector page names Power BI and Power Query Online, not Excel, so Power BI Desktop can open the file directly while Excel needs a CSV. This DuckDB command wrote the 239,088 Texas orders, well under Excel's 1,048,576-row limit, in 0.06 seconds:

Bash
duckdb -c "COPY (SELECT * FROM 'orders_2025.parquet' WHERE state = 'TX') TO 'orders_tx.csv' (HEADER)"

The Parquet Viewer's CSV export does the same job without a terminal.

How to Read a Parquet File in Spark or Databricks

PySpark reads a file, or a whole directory of part files, with one call:

Python
from pyspark.sql import SparkSession, functions as F

spark = SparkSession.builder.appName("read-parquet").getOrCreate()
df = spark.read.parquet("orders_2025.parquet")
df.printSchema()
(df.where(F.col("state") == "TX")
   .groupBy("city")
   .agg(F.count("*").alias("orders"), F.sum("amount_usd").cast("decimal(12,2)").alias("revenue_usd"))
   .orderBy(F.desc("revenue_usd"))
   .show())
text
root
 |-- order_id: long (nullable = true)
 |-- order_date: date (nullable = true)
 |-- customer_name: string (nullable = true)
 |-- city: string (nullable = true)
 |-- state: string (nullable = true)
 |-- zip: string (nullable = true)
 |-- amount_usd: double (nullable = true)

+-----------+------+-----------+
|       city|orders|revenue_usd|
+-----------+------+-----------+
|     Austin| 59982|15141806.86|
|San Antonio| 59812|15126653.56|
|    Houston| 59667|15104170.73|
|     Dallas| 59627|15085591.28|
+-----------+------+-----------+

PySpark 4.2.0 ran it locally on Java 24 in 5.5 seconds, nearly all of it JVM startup. In a Databricks notebook the spark session already exists, so drop the builder line. Databricks' own Parquet guide, updated August 26, 2026, reads from a Unity Catalog volume path such as /Volumes/<catalog>/<schema>/<volume>/ and shows the result with display(df). That guide calls Parquet the most common format for data stored in Databricks, because Delta Lake tables are Parquet files underneath.

In VS Code and R

In VS Code, Microsoft's Data Wrangler extension lists .parquet among the file types it supports: right-click the file in the Explorer and choose Open in Data Wrangler. In R, the arrow package's read_parquet() returns a data frame, and its col_select argument trims columns the way columns= does in pandas. We did not run the R version for this post.

Common Errors

ImportError: Unable to find a usable engine; tried using: 'pyarrow', 'fastparquet'. This is pandas 3.0.6 with neither engine installed. Run pip install pyarrow; pandas tries pyarrow first and only falls back to fastparquet when pyarrow is missing.

pyarrow.lib.ArrowInvalid: Could not open Parquet input source '<Buffer>': Parquet magic bytes not found in footer. Either the file is corrupted or this is not a parquet file. The last 4 bytes are not PAR1. Either the file is not Parquet at all (a CSV renamed to .parquet produced exactly this message), or a download or copy stopped early. Check the tail of the file:

tail -c 4 orders_2025.parquet
PAR1

A truncated copy of the same file failed everywhere, each reader in its own words. DuckDB reported Invalid Input Error: No magic bytes found at end of file 'truncated.parquet'. Polars raised polars.exceptions.ComputeError: parquet: File out of specification: The file must end with PAR1. Spark threw [FAILED_READ_FILE.CANNOT_READ_FILE_FOOTER], with a cause line of is not a Parquet file. Expected magic number at tail, but found [26, 16, 33, 0]. Download the file again rather than trying to repair it, since the footer is the part that is missing.

[PATH_NOT_FOUND] Path does not exist: file:/.../orders_2025.parquet. SQLSTATE: 42K03 comes from Spark when nothing exists at the path. The message prints the full path Spark resolved, with a file: prefix for a local file, so compare it with where the file actually is; an absolute path or a volume path avoids the guesswork.

Unreadable symbols in Notepad++ or a text editor are not an error. Parquet is compressed binary, so open it with one of the readers above instead.

When Not to Do This

Do not convert a Parquet file to CSV just to look at it. The CSV copy here is more than four times larger, and it loses column types, which is how 02108 became 2108. Do not load a multi-gigabyte file whole into pandas either: pass columns, use Polars' scan_parquet, or query it with DuckDB, which reads only the columns a query touches. And before you drop a file with customer data into an online viewer, check that it decodes in the browser rather than uploading the file to a server.

Conclusion

Pick the reader by where the data goes next. For a quick look with nothing installed, the Parquet Viewer showed the schema and first page of the 1,200,000-row orders_2025.parquet in about 4 seconds. In a Python script, pd.read_parquet with a columns list is the shortest path: it read the whole file in 0.02 to 0.09 seconds, against 0.43 to 0.48 seconds for the same data as CSV. When a file is too big to load whole, query it in place with Polars' scan_parquet or DuckDB. Reach for Spark only when the data already lives on a Spark or Databricks cluster, since the local run spent most of its 5.5 seconds starting the JVM. And when Excel is the destination, export a CSV with DuckDB first.

  • Parquet Viewer - open a .parquet file in the browser, check its schema and row groups, search every row and export CSV or JSON without installing Python.
  • Avro Viewer - read Avro container files, the row-oriented format that sits next to Parquet in Kafka and Spark pipelines, with the writer schema from the file header.
  • SQLite Viewer - inspect the tables and rows of a .db file in the browser when a data export arrives as SQLite instead of Parquet.
  • CSV Viewer - page through the CSV you exported from Parquet, with search and filters across every row, before it goes to Excel.

Related Posts

How to Resolve Git Merge Conflicts

Edit the block between the conflict markers, then git add and commit. Command line, VS Code and GitHub, plus why a rebase swaps ours and theirs.

By DevToolLab Team•

How to Enable CORS in Express and Next.js

CORS is enabled on the server, not the browser. Working Express, FastAPI and Next.js configs, how preflight works, and the six errors Chrome prints, with fixes.

By DevToolLab Team•

How to Spot AI Crawlers in Server Logs

Search your access log for GPTBot, ClaudeBot and PerplexityBot, then check each IP against the vendors' published ranges. Includes a Python script and output.

By DevToolLab Team•