Installation & Guide

Getting Started with Bank Statement Parser

Installation instructions for Rust, Python, Docker, and CLI alongside step-by-step extraction quickstarts.

Bank Statement Parser is an open-source, high-throughput financial document engine engineered in Rust with native Python bindings. Follow this guide to install the binary, process multi-format statement documents, or embed the SDK directly into your application stack.

Rust (Cargo)

Crate v0.0.28

Native library integration and standalone CLI compiled directly from source for maximum CPU throughput.

cargo install bankstatementparser

Python (Pip)

PyPI v0.0.28

Pre-compiled PyO3 binary wheels for Linux, macOS (Apple Silicon and Intel), and Windows with optional Parquet extras.

pip install bankstatementparser
# Optional: Apache Parquet and Polars streaming support
pip install "bankstatementparser[parquet]"
pip install "bankstatementparser[polars]"

Docker Engine

OCI Container

Minimal, non-root container image published on GitHub Container Registry for air-gapped CI/CD execution.

docker pull ghcr.io/sebastienrousseau/bankstatementparser:latest

Homebrew Tap

macOS & Linux

Pre-built standalone CLI binary distributed via official Homebrew tap for immediate terminal workstation access.

brew install sebastienrousseau/tap/bankstatementparser
01

Command Line Interface (CLI) Extraction

The standalone binary provides fast parsing for individual bank statement files across digital PDF, CSV, OFX, and QIF formats, emitting standard JSON, CSV, or ISO 20022 CAMT.053 XML:

bash — cli quickstart
# Extract statement PDF to standard JSON
$bankstatementparser --input statement.pdf --output statement.json
 
# Convert OFX banking file to CSV
$bankstatementparser --input statement.ofx --output statement.csv --format csv
 
# Transform PDF statement into ISO 20022 CAMT.053 XML
$bankstatementparser --input statement.pdf --output statement.xml --format camt053

For high-volume transaction processing, enable multi-threaded batch ingestion to parse entire directories in parallel:

bash — batch processing
# Process all statements in a directory using 8 parallel worker threads
$bankstatementparser --batch-dir ./statements/ --output-dir ./extracted_json/ --threads 8
02

Rust Library Integration

Add the dependency to your Cargo.toml:

Cargo.toml
[dependencies]
bankstatementparser = "0.0.28"

Parse any statement into strongly-typed transaction structs with compile-time memory safety:

main.rs
use bankstatementparser::{Parser, StatementFormat};
use std::path::Path;
 
fn main() -> Result<(), Box<dyn std::error::Error>> {
    let parser = Parser::new();
    let statement = parser.parse_file(Path::new("statement.pdf"))?;
 
    println!("Account: {}", statement.account_id);
    println!("Balance: {} {}", statement.closing_balance, statement.currency);
 
    for tx in statement.transactions {
        println!("{} | {} | {}", tx.date, tx.amount, tx.description);
    }
 
    Ok(())
}
03

Python SDK Integration

Extract and validate transaction records within your Python data science and finance workflows with zero native compilation:

python — main.py
from bankstatementparser import (
    BankStatementParser, CamtParser, Bai2Parser, ParquetStreamWriter,
)
 
# 1. Parse statement into strongly-typed Transaction models
parser = BankStatementParser()
statement = parser.parse_file("statement.pdf")
 
# 2. Stream transactions directly to Apache Parquet
with ParquetStreamWriter("transactions.parquet") as writer:
    writer.write_batch(statement.transactions)
 
print(f"Extracted {len(statement.transactions)} transactions into Parquet.")
04

Zero-Telemetry Execution Guarantees

Every extraction pipeline executes exclusively within local process memory. Bank Statement Parser makes zero outbound network calls, transmits no analytical telemetry, and requires no external SaaS API keys:

High-throughput financial document engineering architecture
Air-gapped execution architecture: banking statements are tokenized and validated entirely within local memory with zero external telemetry.