Why build Baruch? Version 0
- non_dev
- narrative
- brandon_thoughts
Backstory
Baruch is not the only tool for building pipelines. It probably is not the only tool to use YAML-esque syntax to do so. So the question is why? Why build this tool when it will take far longer than learning (a likely superior) pre-existing tool? The answer is, of course, I did not know what tools existed. I knew I groaned adjusting pandas code that was built expecting SQL to work with Mongo, I knew I hated vibe coding GUIs for projects at work that I could not quickly verify because I did not recall the syntax for certain operations, and I knew I dreaded that my ETL pipelines were fragile, hard to audit, and were tied explicitly to Python. Worse, I was hand building the ETL code for training and then copying it over into our application code, which meant that I had to make sure that all my training data and live data had identical formats. It was painful and difficult to scale.
Thus, while building an internal tool at work and getting annoyed at the code I was working with Claude to write and working with Docker containers, I thought why not apply the Docker mindset to building datasets? Define a base dataset, write your transforms as layers, and then allow for different data vessels1 to share the data when they share layers, similar to how a Docker container can reuse cached layers from another one. After a discussion with a much wiser engineer (and a consultation with an LLM), I realized that the usefulness of this tool would be limited, the caching mechanic would be very intensive to build and fragile, and I really needed a tool that I could use to describe ETL in every step of the process. So the Baruch project was born, not because other tools did not exist, but because honestly, I did not understand data or code well enough, and I hated consulting with Claude anytime I had to do more complex transforms.
Goals
The goals of Baruch are simple: write a specification for .bml files, build a testing harness to verify that a given implementation is compliant, and implement a Python module that will execute the code. On a more personal development level, the goals are to build a more robust understanding of code and data. For code, specific learning/engineering goals I wish to accomplish:
- Build a parser. Learn about AST, optimization, and all of the fun parts of building a language.
- Work with other modules, libraries, services, etc. I welcome the opportunity to manage the connections between Baruch and, for example, Spark. I hope this will build further skills with these common and useful tools.
- Write good (and hard) code. Professionally, I feel I am getting good at describing and designing code but not at architecting and building it. I want to write matrix operations, do C implementations of critical code, and learn CUDA.
- Learn how to actually write tests!
- Work on an open source project (hopefully with others)
From a data perspective:
- Get comfortable with analytical databases, especially columnar ones. I worked mainly with transactional DBs so far, even when they are not the right option.
- Learn more about querying strategies and other optimizations for data.
What this post is for
Honestly, this post was written late at night before I have typed a single line of actual code. I plan to write an actual post explaining the purpose and niche of Baruch soon. This post is to instead serve as a reminder of my initial naivety and hopefully as a record of where this project started, so when I return for version 1 of this post, I can see how we have adjusted.
Current (Very Rough) Roadmap (Likely too aggressive)
This is aspirational and will likely not track reality, particularly on dates. Still, I hope to follow it, but a live roadmap is available in the repo.
By End of September 2026 V 0.2.x
- Initial Specification Fully Written
- Details on sources, common transformations (resample, pivot, ffill/bfill, etc), outputs and connections between nodes.
- Unit tests for each node
- Parser written
- Implementation of sources (Mongo, ANSI SQL, flat files)
End of November 2026 V 0.3.x
- Expand Specification to include other transformations, error handling, and improve harness
- Integration tests.
- Specify how to deal with different errors
- Implement outputs
End of December 2026 V 0.4.x
- Implement common transformations
End of February 2027 V 0.5.x
- Implement full specifications
- Add other common sources and outputs
End of April 2027 V 1.0.0
- System testing
- End to end testing
- Complete documentation of code
Footnotes
-
The term I am using as analog for a Docker container in my data-Docker dream. ↩