Multi-Format Dialect Engine
Deterministic state machines parse digital and scanned PDFs, irregular CSV layouts, SGML OFX/QFX, BAI2 cash management, and SWIFT MT940/MT942 messages without brittle regex or external cloud OCR.
Financial Document Engineering · Open Source
An open-source, high-throughput financial document parsing engine engineered in Rust with native Python bindings. Converts CAMT, PAIN.001, MT940, MT942, OFX/QFX, BAI2, CSV, and PDF bank statements into validated JSON, Apache Parquet, and ISO 20022 transaction streams entirely on your own infrastructure.
Developer Quickstart
Deploy as a standalone CLI tool, embed as a high-assurance Rust crate in your microservices, or integrate into Python data pipelines with zero external runtime overhead.
Engine Architecture
Industrial-grade financial parsing combining memory-safe Rust execution, native Python bindings, and dual Apache/MIT licensing.
Deterministic state machines parse digital and scanned PDFs, irregular CSV layouts, SGML OFX/QFX, BAI2 cash management, and SWIFT MT940/MT942 messages without brittle regex or external cloud OCR.
Strict mathematical balance validation ensures opening balance plus net credits and debits precisely equals closing balance to the penny before record emission.
Normalize disparate bank statements directly into canonical ISO 20022 CAMT.053 XML envelopes, Apache Parquet streams, or strongly-typed JSON schemas ready for ERP, GL, and accounting ingestion.
The Processing Pipeline
Stream PDF bytes, CSV feeds, BAI2 files, or SWIFT MT940/MT942 messages into memory-safe zero-copy buffers.
Sub-millisecond stream loadingDeterministic lexical analysis identifies account metadata, booking dates, line items, and structured :86: sub-tags.
Deterministic layout analysisArithmetic balance reconciliation confirms credit/debit integrity, checks PDF forensics, and validates IBAN checksums.
Mod-97 and balance proofsOutput strongly-typed JSON, streaming Apache Parquet, Plain-Text Ledgers (hledger, beancount), or ISO 20022 messages.
Ready for ERP, Lakehouse & GL
Corporate Treasury Automation
Eliminate manual statement keying, error-prone spreadsheets, and expensive cloud OCR APIs. Normalize multi-bank statements directly into your general ledger, ERP, and treasury management workflows with sub-second turnaround.
View Treasury Playbook
High-Performance Pipeline
Engineered for high-frequency financial platforms requiring bounded latency, zero-allocation loops, and strict ISO 20022 data models.
View Architecture BlueprintZero-Telemetry Privacy
Bank Statement Parser is 100% self-contained and operates entirely on your local machine, private VPC, or air-gapped on-premise servers. Zero analytics, zero telemetry, and zero outbound network calls—guaranteeing complete GDPR, GLBA, and banking confidentiality compliance.
Clear Answers
No. Bank Statement Parser is 100% self-contained and operates entirely on your local machine or private cloud server. It contains zero analytics, zero telemetry, and zero outbound network calls, ensuring complete compliance with GDPR, HIPAA, GLBA, and banking confidentiality regulations.
Bank Statement Parser natively supports eight primary banking formats: ISO 20022 CAMT (camt.052, camt.053, camt.054), ISO 20022 PAIN.001 (Credit Transfer), SWIFT MT940 and MT942 messages, BAI2 (Bank Administration Institute), Open Financial Exchange (OFX 1.x/2.x and QFX), CSV files (with automatic delimiter and date detection), and digital/scanned PDF statements.
For digital and vector PDFs, the engine extracts structured text streams directly with zero loss. For scanned image statements, an optional local OCR module performs deterministic optical layout parsing with zero external cloud API dependencies.
Yes. Bank Statement Parser is dual-licensed under the Apache-2.0 and MIT open-source licenses. You can freely integrate it into commercial SaaS products, enterprise backends, and internal financial pipelines.
Developer Quickstart
Get started in minutes with the Rust CLI or Python package across Linux, macOS, and Windows.
Install Bank Statement Parser