Apache Livy 客户端
项目描述
生活
Apache Livy 客户端
安装库
pip install livyc
导入库
from livyc import livyc
设置 livy 配置
data_livy = {
"livy_server_url": "localhost",
"port": "8998",
"jars": ["org.postgresql:postgresql:42.3.1"]
}
让我们尝试向 Apache Livy Server 启动一个 pySpark 脚本
params = {"host": "localhost", "port":"5432", "database": "db", "table":"staging", "user": "postgres", "password": "pg12345"}
pyspark_script = """
from pyspark.sql.functions import udf, col, explode
from pyspark.sql.types import StructType, StructField, IntegerType, StringType, ArrayType
from pyspark.sql import Row
from pyspark.sql import SparkSession
df = spark.read.format("jdbc") \
.option("url", "jdbc:postgresql://{host}:{port}/{database}") \
.option("driver", "org.postgresql.Driver") \
.option("dbtable", "{table}") \
.option("user", "{user}") \
.option("password", "{password}") \
.load()
n_rows = df.count()
spark.stop()
"""
创建一个 livyc 对象
lvy = livyc.LivyC(data_livy)
创建到 Apache Livy 服务器的新会话
session = lvy.create_session()
在 Apache Livy 服务器中发送和执行脚本
lvy.run_script(session, pyspark_script.format(**params))
访问会话中可用的变量“n_rows”
lvy.read_variable(session, "n_rows")
贡献和反馈
有关此存储库的任何想法或反馈?帮助我改进它。
作者
- 由拉美西斯·亚历山大·科拉斯佩·瓦尔迪兹创作
- 创建于 2022 年
执照
该项目根据 MIT 许可条款获得许可。
项目详情
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
livyc-0.0.14.tar.gz
(19.2 kB
view hashes)
Built Distribution
livyc-0.0.14-py3-none-any.whl
(5.7 kB
查看哈希)