Selaa lähdekoodia

Merge pull request #39 from lnyemba/v2.2.0

V2.2.0 - merge
Steve L. Nyemba 1 viikko sitten
vanhempi
commit
12eba9d182
5 muutettua tiedostoa jossa 67 lisäystä ja 42 poistoa
  1. 31 33
      README.md
  2. 1 1
      transport/info.py
  3. 1 1
      transport/other/files.py
  4. 24 6
      transport/other/http.py
  5. 10 1
      transport/sql/duckdb.py

+ 31 - 33
README.md

@@ -1,44 +1,42 @@
-# Introduction
+# Data-Transport
 
-This project implements an abstraction of objects that can have access to a variety of data stores, implementing read/write with a simple and expressive interface. This abstraction works with **NoSQL**, **SQL** and **Cloud** data stores and leverages **pandas**.
+A powerful abstraction layer for seamless data communication across diverse systems. **data-transport** allows you to interact with NoSQL, SQL, and Cloud storage using a consistent interface powered by **Pandas** and **SQLAlchemy**.
 
-# Why Use Data-Transport ?
+## Why Choose data-transport?
 
-Data transport is a simple framework that enables read/write to multiple databases or technologies that can hold data. In using **data-transport**, you are able to:
+*   **Unified Interface:** Connect to PostgreSQL, MySQL, MongoDB, S3, etc., using the same consistent code.
+*   **Security First:** Prevents the dissipation of database connectivity information to protect against security breaches.
+*   **Simplicity & Power:** Leverages Pandas DataFrames and SQLAlchemy for intuitive data manipulation.
+*   **Robust Pipelines:** Easily integrate pre-processing and post-processing as unified pipelines.
+*   **CLI Integration:** Includes a dedicated CLI for registry management and ETL task execution.
 
-- Enjoy the simplicity of **data-transport** because it leverages SQLAlchemy & Pandas data-frames.
-- Share notebooks and code without having to disclosing database credentials. 
-- Seamlessly and consistently access to multiple database technologies at no cost
-- No need to worry about accidental writes to a database leading to inconsistent data
-- Implement consistent pre and post processing as a pipeline i.e aggregation of functions
-- **data-transport** is open-source under MIT License https://github.com/lnyemba/data-transport
+## Supported Features
 
-## Installation
-
-Within the virtual environment perform the following, the options for installation are:
-
-**sql**       - by default postgresql, mysql, sqlserver, sqlite3+, duckdb 
-
-    pip install data-transport[cloud,nosql,other,all]git+https://github.com/lnyemba/data-transport.git
-
-Options to install components in square brackets, these components are 
+| Component | Technologies Covered |
+| :--- | :--- |
+| **SQL** | PostgreSQL, MySQL, SQL Server, SQLite3+, DuckDB |
+| **NoSQL** | MongoDB, CouchDB |
+| **Warehouse** | Apache Iceberg, Apache Drill |
+| **Cloud** | Nextcloud, S3 |
+| **Other** | Files, RabbitMQ, HTTP |
 
-**warehouse** - Apache Iceberg, Apache Drill
-
-**cloud**  - to support nextcloud, s3
-
-**nosql** - support for mongodb, couchdb
-
-**other**  - support for files, rabbitmq, http
+## Installation
 
-    pip install data-transport[nosql,cloud,warehouse,all]@git+https://github.com/lnyemba/data-transport.git
+Install the core package and your desired components:
 
-## Additional features
+```bash
+# Basic installation with default SQL support
+pip install data-transport@git+https://github.com/lnyemba/data-transport
 
-    - In addition to read/write, there is support for functions for pre/post processing
-    - CLI interface to add to registry, run ETL
-    - scales and integrates into shared environments like apache zeppelin; jupyterhub; SageMaker; ...
+# Full suite (SQL, NoSQL, Cloud, Warehouse)
+pip install "data-transport[nosql,cloud,warehouse,all]"@git+https://github.com/lnyemba/data-transport.git
+```
 
-## Learn More
+## Advanced Capabilities
+*   **Automated Pipelines:** Seamlessly aggregate functions for data cleaning and transformation.
+*   **Portability:** Share notebooks and scripts without exposing raw credentials.
+*   **Scalable Integration:** Compatible with environments like Apache Zeppelin, JupyterHub, and SageMaker.
 
-We have available notebooks with sample code to read/write against mongodb, couchdb, Netezza, PostgreSQL, Google Bigquery, Databricks, Microsoft SQL Server, MySQL ... Visit [data-transport homepage](https://healthcareio.the-phi.com/data-transport)
+---
+[Learn More at the Project Website](https://healthcareio.the-phi.com/data-transport)
+License: [MIT](https://github.com/lnyemba/data-transport)

+ 1 - 1
transport/info.py

@@ -1,6 +1,6 @@
 __app_name__  = 'data-transport'
 __author__ = 'The Phi Technology'
-__version__= '2.2.42'
+__version__= '2.2.46'
 __email__  = "info@the-phi.com"
 __edition__= 'community'
 __license__=f"""

+ 1 - 1
transport/other/files.py

@@ -34,7 +34,7 @@ class Reader (File):
 	def read(self,**args):
 		_path = self.path if 'path' not in args else args['path']
 		_delimiter = self.delimiter if 'delimiter' not in args else args['delimiter']
-		_df =  pd.read_csv(_path,delimiter=self.delimiter)
+		_df =  pd.read_csv(_path,delimiter=_delimiter)
 		if 'query' in args :
 			_query = args['query']
 			_df = _df.query(_query)

+ 24 - 6
transport/other/http.py

@@ -19,12 +19,14 @@ class Reader:
 	def __init__(self,**_args):
 		self._url = _args['url']
 		self._headers = None if 'headers' not in _args else _args['headers']
+		self._data = None if 'data' not in _args else _args['data']
+		self._method = 'get' if 'method' not in _args else _args['method'].lower()
 					
 	# def isready(self):
 	# 	return self.file_length > 0
 	def format(self,_response):
 		_mimetype= _response.headers['Content-Type']
-		if _mimetype == 'text/csv' or 'text/csv':
+		if 'text/plain' in _mimetype or 'text/csv' in _mimetype:
 			_content = _response.text
 			return pd.read_csv(StringIO(_content))
 		#
@@ -32,12 +34,28 @@ class Reader:
 		#
 		
 		return _response.text
+	def get (self,key,_args):
+		"""
+		This function inspects an argument and tries to determine if the corresponding attribute is set 
+			i.e the attribute will be prefixed by underscore
+		
+		"""
+
+		_attr = f'_{key}'
+		return _args[key] if key in _args else (getattr(self,_attr) if hasattr(self,_attr) else None)
 	def read(self,**_args):
-		if self._headers :
-			r = requests.get(self._url,headers = self._headers)
-		else:
-			r = requests.get(self._url,headers = self._headers)
-		return self.format(r)
+		_method  	= self.get('method',_args)
+		_headers	= self.get('headers',_args)
+		_data 		= self.get('data',_args)
+		_url		= self.get('url',_args)
+		_requestPpointer = getattr(requests,_method)
+		_resp =  _requestPpointer(_url,headers=_headers,data=_data)
+		return self.format(_resp)
+		# if self._headers :
+		# 	r = requests.get(self._url,headers = self._headers)
+		# else:
+		# 	r = requests.get(self._url,headers = self._headers)
+		# return self.format(r)
 		
 class Writer:
 	"""

+ 10 - 1
transport/sql/duckdb.py

@@ -20,7 +20,16 @@ class Reader(Duck,BaseReader) :
         Duck.__init__(self,**_args)
         BaseReader.__init__(self,**_args)
     def _get_uri(self,**_args):
-        return super()._get_uri(**_args),{'connect_args':{'read_only':True}}
+        #
+        # if we are working with an in-memory database we can NOT set the attributes to be read-only
+        # something to do with SQL-Alchemy
+        p = self.database.strip().startswith(":") and self.database.strip().endswith(":")
+        if not p :
+            return super()._get_uri(**_args),{'connect_args':{'read_only':True}}
+        else:
+            #
+            # we have an in-memory database 
+            return super()._get_uri(**_args),{}
 class Writer(Duck,BaseWriter):
     def __init__(self,**_args):
         Duck.__init__(self,**_args)