Software Engineer - Data Engineering
Make mosaicod interoperable with the modern data ecosystem (Delta Lake, Iceberg, Arrow, and DataFusion), building the conversion layer in Rust, open source.
Mosaico is an open-source data platform for robotics and physical AI. We’re solving a hard infrastructure problem: the tools that exist today for storing, indexing, and retrieving high-frequency sensor data are either too primitive or too painful to work with at scale.
We’re building the missing piece, and doing it in the open. Apache 2.0 licensed, public roadmap, real community. We care about the codebase being something people actually want to read and contribute to, not just use. In Italy, companies that operate this way are still a rarity, and we think that’s exactly what makes this interesting.
We’re based in Reggio Emilia, working with teams in defence, automotive, and agritech, and we’re building a commercial offering on top of the open-source core.
The role
Mosaico already uses Apache Arrow, DataFusion, and Parquet internally as core parts of its storage and query engine. The next step is making the platform interoperable with the broader data engineering ecosystem: Delta Lake, Apache Iceberg, and the other open table formats that modern ETL pipelines are built around.
You’ll design and implement the conversion and integration layer that makes this possible, both reading data into Mosaico from external systems and exporting Mosaico data into formats that downstream pipelines can consume natively. This means working deep inside mosaicod in Rust, with a clear understanding of how these formats work under the hood and where the interesting tradeoffs are.
What you’ll work on
- Design and implement read and write conversion layers between mosaicod and open table formats such as Delta Lake and Apache Iceberg
- Extend the existing DataFusion and Parquet integration to support new formats and interoperability patterns
- Define the API surface that exposes interoperability features to SDK clients and downstream pipelines
- Contribute to the open-source codebase and engage with the broader Arrow and DataFusion communities
- Participate in architectural decisions about how mosaicod evolves as a node in larger data ecosystems
What we’re looking for
- Strong Rust, production-grade and used in a systems or data engineering context
- Deep knowledge of Apache Arrow and DataFusion, not just as a user but with an understanding of the internals
- Hands-on experience with open table formats: Delta Lake, Apache Iceberg, Apache Hudi, or similar
- Solid understanding of how modern ETL pipelines are built and how data moves between systems
- Comfortable working at the intersection of systems engineering and data engineering
- Autonomous and able to drive technical decisions in a small team
- Fluent in English, written and spoken
Nice to have
- Experience contributing to Arrow, DataFusion, or related open-source projects
- Familiarity with the robotics or sensor data domain
- Background in query engine internals or storage format design
- Experience with cloud object storage and its interaction with open table formats
What we offer
- Deep technical ownership of the interoperability layer inside mosaicod
- A codebase that already uses Arrow and DataFusion seriously, so you’re not starting from zero
- Work that ships as open source under Apache 2.0 and powers a commercial product
- Flexible hours, full remote with optional office access in Reggio Emilia
- Competitive salary, discussed openly based on your level and experience
- Welfare package
- Stock options
Salary range
€70K – 100K, adjusted based on location, experience, and level.
Contacts
Interested? Write to [email protected]