Git for data that turns object storage into versioned, branchable data lake repositories
Apache-2.0
- Go
- JavaScript
- Java

About lakeFS
lakeFS is data version control for data lakes - Git for data. It turns object storage into a Git-like repository, so teams can branch, commit, and roll back data the same way they manage code, and run repeatable, atomic data lake operations across ETL, analytics, and data science.
Branching gives every job a full, zero-copy snapshot of production data for isolated dev and test environments. Hooks enforce write-audit-publish gates that validate schema, format, or governance rules before data is merged, and time travel lets you reproduce or roll back to any past version after a bad write.
lakeFS runs on AWS S3, Azure Blob Storage, and Google Cloud Storage. It is API compatible with S3 and works with Spark, Hive, AWS Athena, DuckDB, and Presto.
Key features
- Zero-copy branches over production data lake storage
- Hooks that enforce write-audit-publish gates before merge
- Time travel and rollback to any past version
- S3-compatible API over object storage
- Works with Spark, Hive, Athena, DuckDB, and Presto
Details
- First released
- 2019
- Type
- Data version control
- Language
- Go
- Storage
- AWS S3 · Azure Blob · Google Cloud Storage
- Compatibility
- S3 API compatible
- Maintainer
- Treeverse
