Ingen beskrivning

Steve Nyemba de4ee2fcfa bug fixes ... windows runner, files 11 månader sedan
bin de4ee2fcfa bug fixes ... windows runner, files 11 månader sedan
info 66d881fdda upgrade pyproject.toml, bug fix with registry 1 år sedan
notebooks 1a8112f152 adding iceberg notebook 1 år sedan
transport a4597d4a8c adding queries to files 11 månader sedan
.gitignore 92bf0600c3 .. 2 år sedan
README.md 4c2efc2892 documentation ... readme 1 år sedan
pyproject.toml de4ee2fcfa bug fixes ... windows runner, files 11 månader sedan
requirements.txt 8d4ecd7a9f S3 Requirments file 9 år sedan

README.md

Introduction

This project implements an abstraction of objects that can have access to a variety of data stores, implementing read/write with a simple and expressive interface. This abstraction works with NoSQL, SQL and Cloud data stores and leverages pandas.

Why Use Data-Transport ?

Data transport is a simple framework that:

  • easy to install & modify (open-source)
  • enables access to multiple database technologies (pandas, SQLAlchemy)
  • enables notebook sharing without exposing database credential.
  • supports pre/post processing specifications (pipeline)

Installation

Within the virtual environment perform the following :

pip install git+https://github.com/lnyemba/data-transport.git

Options to install components in square brackets

pip install data-transport[nosql,cloud,warehouse,all]@git+https://github.com/lnyemba/data-transport.git

Additional features

- In addition to read/write, there is support for functions for pre/post processing
- CLI interface to add to registry, run ETL
- scales and integrates into shared environments like apache zeppelin; jupyterhub; SageMaker; ...

Learn More

We have available notebooks with sample code to read/write against mongodb, couchdb, Netezza, PostgreSQL, Google Bigquery, Databricks, Microsoft SQL Server, MySQL ... Visit data-transport homepage