# Creating a separate database for datasets created by Python pipelines

**URL:** <https://discourse.getdbt.com/t/creating-a-separate-database-for-datasets-created-by-python-pipelines/4891>\
**Category:** In-Depth Discussions\
**Tags:** staging, lineage\
**Created:** [September 8, 2022, 2:56pm UTC](https://discourse.getdbt.com/t/creating-a-separate-database-for-datasets-created-by-python-pipelines/4891 "2022-09-08T14:56:45Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![nancy.chelaru](https://avatars.discourse-cdn.com/v4/letter/n/e19adc/32.png) [@nancy.chelaru](https://discourse.getdbt.com/u/nancy.chelaru)\
**Post date:** [September 8, 2022, 2:56pm UTC](https://discourse.getdbt.com/t/creating-a-separate-database-for-datasets-created-by-python-pipelines/4891/1 "2022-09-08T14:56:45Z")

</div>

Hello!

Please let me know if this is the wrong category for post this!

A bit of background on our current database set up. We have a staging database for preparing datasets from external provider databases, a development database that mirrors the production database. We also have a number of Python pipelines deployed on a server producing datasets, such as forecasting results and web scraping data.

We are currently trying to create a proper staging to data mart data flow. Should we 1) create a separate database for these pipeline-produced datasets that then are brought into the staging database, treating them as a raw data source or 2) put them into a “pipelines” schema directly in the staging database because we have control over the formats and quality of these datasets (this is referencing a [post](https://discourse.getdbt.com/t/how-we-structure-our-dbt-projects/355/11) by Claire about putting seeds directly in the staging database since they can already be put into a “staging” format).

Any thoughts would be greatly appreciated! 🙂

Nancy

---

<div class="post-metadata">

**Author:** ![joellabes](https://sea2.discourse-cdn.com/flex020/user_avatar/discourse.getdbt.com/joellabes/32/3980_2.png) [@joellabes](https://discourse.getdbt.com/u/joellabes)\
**Post date:** [September 8, 2022, 8:03pm UTC](https://discourse.getdbt.com/t/creating-a-separate-database-for-datasets-created-by-python-pipelines/4891/2 "2022-09-08T20:03:43Z")

</div>

Hey @nancy.chelaru, this is a really interesting question! I think my answer depends on the specifics of

> [@nancy.chelaru](#):
>
> we have control over the formats and quality of these datasets.

How much transformation and cleaning is happening in Python before they’re being loaded to the database, and does your team control that process directly? If the data is already in a perfect format for analysis, and that format isn’t going to change ~~ever~~ for a long time, you could probably get away with loading it directly into a `staging` db.

“Could probably get away with” is not exactly inspiring though! On the whole, I’d recommend still loading it into a separate `raw` database. Your staging layer might wind up being `select * from "raw"."pipelines"."web"` for now, but it maintains optionality if your data format changes in the future - you can make changes in your staging layer to maintain the interface/contract that your marts rely on.

---

<div class="post-metadata">

**Author:** ![nancy.chelaru](https://avatars.discourse-cdn.com/v4/letter/n/e19adc/32.png) [@nancy.chelaru](https://discourse.getdbt.com/u/nancy.chelaru)\
**Post date:** [September 8, 2022, 8:31pm UTC](https://discourse.getdbt.com/t/creating-a-separate-database-for-datasets-created-by-python-pipelines/4891/3 "2022-09-08T20:31:48Z")

</div>

Thank you so much for your reply, @joellabes !

Great to hear your take on it! Our team does have complete control over these pipelines. I too was leaning towards having a separate `raw` database for the resultant datasets. It just didn’t quite feel right for them to land directly in the `staging` database, which is supposed to, as you said, a layer to maintain data contracts that downstream models depend on. Having them in a separate database also makes the lineage graph clearer.
