> ## Documentation Index
> Fetch the complete documentation index at: https://docs.0xarchive.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Bulk Exports quickstart

> Buy one day of Hyperliquid BTC trades in the Data Catalog, download the Parquet file, and read its rows in Python with pandas, Polars, or DuckDB.

Buy one UTC day of Hyperliquid BTC trades, download the file, and read it in Python. The code reads any Hyperliquid perpetual trades file, so you can choose another perpetual market or a longer range instead. Set `FILE` to the name of the file you download.

The ordering and download steps are the same for every data type. Other data types have different columns, so the code in steps 3 to 5 does not apply to them; their layouts are on [File columns](/schemas/export-files).

You need a 0xArchive account; you can build the order signed out and sign in at checkout. Orders have a `$10` minimum; the quote shows the price for the day you pick.

## 1. Order the file

<Steps>
  <Step title="Choose the market and data type">
    Open the [Data Catalog](https://0xarchive.io/data), search for `BTC`, and open the Hyperliquid perpetual market's catalog page. Under **Select data schemas**, keep **Trades** selected and clear the others.
  </Step>

  <Step title="Choose one UTC day">
    Under **Select date range**, choose **Custom** and pick the same day as start and end, for example 2026-10-01. A day runs from 00:00 to 24:00 UTC. Pick a day before today so the file holds the whole day.
  </Step>

  <Step title="Add it to the cart">
    The **Add to Cart** button shows the price for your selection. Add the line, then select **Review order** in the cart.
  </Step>

  <Step title="Pay or use export credits">
    The **Review export order** page shows the order and its **Total today**, after export credits and any discount. Select **Confirm & Pay**. If you are signed out, the button reads **Sign in to continue** and brings you back to the order afterwards. When export credits cover the total, there is no payment step; otherwise you pay by card in Stripe Checkout.
  </Step>

  <Step title="Wait for Ready to Download">
    You land on the order page in your dashboard, titled **Export Order**, at `/dashboard/exports/<order ID>`. It shows **Queued**, then **Processing** with a percentage, then **Ready to Download**. You also get an email when the files are ready.
  </Step>
</Steps>

## 2. Download the files

Under **Download Files**, the order page lists two files:

* `BTC_trades_2026-10-01_to_2026-10-01.parquet`, the trades.
* `README.md`, which lists the files and describes every column with its type and meaning.

Select a file name to download it. To download on a server instead, copy the file's link and fetch it with curl. Keep the quotes: the link carries a signature in its query string.

```bash theme={"theme":"github-dark"}
curl -o BTC_trades_2026-10-01_to_2026-10-01.parquet "<download link>"
```

Download links last 7 days. Open the order page again for fresh links; downloads are available for 30 days after the export completes.

## 3. Read the rows

Pick one library and use it for the rest of this guide. Each tab installs what it needs and sets `FILE` to the name of the file you downloaded; change it if you chose another market or range. The later steps reuse it in the same Python session.

<CodeGroup>
  ```python pandas theme={"theme":"github-dark"}
  # pip install pandas pyarrow
  import pandas as pd

  FILE = "BTC_trades_2026-10-01_to_2026-10-01.parquet"

  df = pd.read_parquet(FILE)
  print(df.columns.tolist())
  print(df[["timestamp", "side", "price", "size", "crossed"]].head())
  ```

  ```python Polars theme={"theme":"github-dark"}
  # pip install polars
  import polars as pl

  FILE = "BTC_trades_2026-10-01_to_2026-10-01.parquet"

  df = pl.read_parquet(FILE)
  print(df.columns)
  print(df.select("timestamp", "side", "price", "size", "crossed").head())
  ```

  ```python DuckDB theme={"theme":"github-dark"}
  # pip install duckdb
  import duckdb

  FILE = "BTC_trades_2026-10-01_to_2026-10-01.parquet"

  duckdb.sql(f"""
      SELECT timestamp, side, price, size, crossed
      FROM '{FILE}'
      ORDER BY timestamp
      LIMIT 5
  """).show()
  ```
</CodeGroup>

With pandas, the output looks like this. It is illustrative: the columns are the layout of a Hyperliquid perpetual trades file, and the rows are real BTC fills from 2026-10-01 written in that layout. Your rows depend on the market and day you buy.

```text theme={"theme":"github-dark"}
['timestamp', 'coin', 'side', 'price', 'size', 'trade_id', 'order_id', 'user_address', 'crossed', 'direction', 'fee', 'closed_pnl', 'start_position', 'tx_hash', 'builder_address', 'builder_fee', 'deployer_fee', 'priority_gas', 'cloid', 'twap_id']
                         timestamp side    price     size  crossed
0 2026-10-01 00:00:00.530000+00:00    B  83607.0  0.00100     True
1 2026-10-01 00:00:00.530000+00:00    S  83607.0  0.00100    False
2 2026-10-01 00:00:00.597000+00:00    S  83606.0  0.00020     True
3 2026-10-01 00:00:00.597000+00:00    B  83606.0  0.00020    False
4 2026-10-01 00:00:01.442000+00:00    S  83607.0  0.01377    False
```

How to read it:

* Each row is one fill. A trade appears twice, once for each side, and both rows share the same `trade_id`.
* `side` is `B` for a buy and `S` for a sell, from the point of view of the row's account. Export files write `S` where the REST API and WebSocket return `A`, so map `S` to `A` before joining this file to API data.
* `crossed` is `True` on the taker's row.
* `timestamp` is UTC with millisecond precision. `price` and `size` are floating-point numbers.

[File columns](/schemas/export-files) describes every column of every file.

## 4. Count each trade once

Keep the taker rows to count each trade once. This sums the day's taker buy and sell size, in the same session as step 3:

<CodeGroup>
  ```python pandas theme={"theme":"github-dark"}
  takers = df[df["crossed"]]
  print(takers.groupby("side")["size"].sum())
  ```

  ```python Polars theme={"theme":"github-dark"}
  print(df.filter(pl.col("crossed")).group_by("side").agg(pl.col("size").sum()).sort("side"))
  ```

  ```python DuckDB theme={"theme":"github-dark"}
  duckdb.sql(f"""
      SELECT side, sum(size) AS size
      FROM '{FILE}'
      WHERE crossed
      GROUP BY side
      ORDER BY side
  """).show()
  ```
</CodeGroup>

## 5. Check what you bought

Every Parquet file carries a description of itself, so a file keeps its context after it leaves the download folder. Read it in the same session:

<CodeGroup>
  ```python pandas theme={"theme":"github-dark"}
  import pyarrow.parquet as pq  # installed with pandas above

  meta = pq.read_schema(FILE).metadata
  for key in [b"exchange", b"symbol", b"data_type", b"date_range"]:
      print(key.decode(), "=", meta[key].decode())
  ```

  ```python Polars theme={"theme":"github-dark"}
  meta = pl.read_parquet_metadata(FILE)
  for key in ["exchange", "symbol", "data_type", "date_range"]:
      print(key, "=", meta[key])
  ```

  ```python DuckDB theme={"theme":"github-dark"}
  duckdb.sql(f"""
      SELECT decode(key) AS key, decode(value) AS value
      FROM parquet_kv_metadata('{FILE}')
      WHERE decode(key) IN ('exchange', 'symbol', 'data_type', 'date_range')
  """).show()
  ```
</CodeGroup>

For this order, the pandas and Polars versions print the following. It is illustrative, built from the metadata fields the export service writes:

```text theme={"theme":"github-dark"}
exchange = hyperliquid
symbol = BTC
data_type = trades
date_range = 2026-10-01 to 2026-10-01
```

The same metadata holds the export time (`exported_at`), the order ID (`export_job_id`), the license, and a `column.<name>.description` entry for each column.

## Large files

DuckDB reads only the columns a query uses and streams through the file, so it can aggregate files larger than memory without loading them. In the same session, this computes taker size and the volume-weighted price for each hour:

```python theme={"theme":"github-dark"}
# pip install duckdb
import duckdb

duckdb.sql(f"""
    SELECT date_trunc('hour', timestamp) AS hour,
           sum(size) AS taker_size,
           sum(price * size) / sum(size) AS vwap
    FROM '{FILE}'
    WHERE crossed
    GROUP BY hour
    ORDER BY hour
""").show()
```

To read several files at once, replace `'{FILE}'` with a pattern such as `'BTC_trades_*.parquet'`.

## If something goes wrong

* **The quote says no data is available.** The market has no rows of that data type on those dates. Choose dates inside the range the Data Catalog shows for the data type.
* **You left Stripe Checkout without paying.** The order stays on its order page as **Awaiting Payment**. Open the page again later to resume payment. Unpaid orders expire, so if the page no longer offers payment, build the order again in the Data Catalog.
* **The order page shows Unavailable.** No file was produced, for one of two reasons:
  * **You never paid**, so the order expired. Build it again in the Data Catalog.
  * **You paid and the export failed.** Export credits it used return to your balance automatically; for a card payment, [contact support](https://0xarchive.io/contact).

## Next step

<Card title="Choose data types for your own dataset" icon="table" href="/export-schemas">
  See which data types each venue offers, what each file contains, and how prices are worked out.
</Card>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.