
在 YARN 集群部署 Jupyter Enterprise Gateway构建企业级 Spark 数据科学平台【免费下载链接】enterprise_gatewayA lightweight, multi-tenant, scalable and secure gateway that enables Jupyter Notebooks to share resources across distributed clusters such as Apache Spark, Kubernetes and others.项目地址: https://gitcode.com/gh_mirrors/en/enterprise_gateway当你的数据团队规模扩大、Spark 任务越来越重时传统的单机 Jupyter 架构很快会成为瓶颈。Jupyter Enterprise Gateway正是为解决这一问题而生的轻量级、多租户、可扩展的安全网关它让 Jupyter Notebook 能够把内核作为受管理资源启动到分布式集群中在YARN 集群部署场景下与 Apache Spark 无缝集成。本文将用最通俗的方式带你一步步完成Jupyter Enterprise Gateway YARN 集群部署构建真正的企业级 Spark 数据科学平台。为什么需要 Jupyter Enterprise Gateway先看一张对比动图直观感受部署前后的巨大差异在传统架构中所有 Jupyter 内核和 Spark Driver 都以本地进程方式运行在一台服务器上。单机内存、CPU 有限内核数量增长极慢资源严重浪费同时所有用户共享同一权限安全性也难以保障。部署 Jupyter Enterprise Gateway 后内核作为受管理资源运行在集群所有节点上随节点数增加可同时运行的内核数量几乎线性增长真正发挥了集群的计算潜力。Jupyter Enterprise Gateway 在 YARN 上的工作原理Jupyter Enterprise Gateway 在 YARN 集群中的角色可以理解为智能调度中枢它接收客户端的 Notebook 请求通过内置的 YARN 进程代理Process Proxy把内核以 Spark 应用的形式提交到 YARN 资源管理器内核的 Driver 和 Executor 都由 YARN 统一调度。更具体的架构关系如下图所示用户通过 HTTPS 连接到网关节点网关将 Spark 内核提交到 YARN 集群Driver/Executor 在 YARN 容器中运行并通过用户模拟Impersonation机制保障多租户安全。如果你使用的是 HDPHortonworks Data Platform发行版还可通过 Knox 反向代理进一步增强企业级安全具体架构可参考下图部署前的准备工作清单开始部署前请确认以下条件项目要求操作系统支持 Linux推荐 CentOS / UbuntuPython 环境建议使用 AnacondaPython 3.11Spark 集群已部署 Hadoop YARN Spark记录SPARK_HOME路径节点要求内核包需安装到集群的每个可用节点Scala 内核除外 提示官方完整步骤见 deploy-yarn-cluster.md 与 installing-eg.md。一键安装步骤快速完成 Jupyter Enterprise Gateway 安装在 YARN 集群主节点上推荐位置用 pip 或 conda 即可完成安装# 方式一使用 pip 从 PyPI 安装 pip install --upgrade jupyter_enterprise_gateway # 方式二使用 conda 从 conda-forge 安装 conda install -c conda-forge jupyter_enterprise_gateway安装完成后可以用jupyter enterprisegateway --help-all查看全部配置项用默认参数启动jupyter enterprisegateway --ip0.0.0.0 --port_retries0配置 YARN 集群模式内核Python / Scala / R 三选一Jupyter Enterprise Gateway 开箱即用地支持三种内核在 YARN 集群模式下各有对应的示例内核规格kernelspecPythonspark_python_yarn_clusterIPython 内核Scalaspark_scala_yarn_clusterApache Toree 内核Rspark_R_yarn_clusterIRkernel 内核将对应内核规格解压到内核目录后其kernel.json中通过process_proxy指定了 YARN 进程代理类例如 Python 内核使用 yarn.py 中的YarnClusterProcessProxy并通过SPARK_OPTS指定--master yarn --deploy-mode cluster提交方式。项目中的示例配置见 spark_python_yarn_cluster/kernel.json。 提示示例内核规格中同时包含 client 模式如spark_python_yarn_client两种模式均可使用集群模式更利于资源隔离。关键环境变量配置让网关与 YARN 顺利对接要让 Jupyter Enterprise Gateway 与 YARN 集群协同工作必须正确设置以下环境变量环境变量作用示例SPARK_HOME指向 Spark 安装路径/usr/hdp/current/spark2-clientEG_YARN_ENDPOINTYARN 资源管理器地址网关与集群不在同一网络时必填http://rm-node:8088/ws/v1/clusterEG_ALT_YARN_ENDPOINTYARN 高可用时的备用资源管理器地址http://rm-standby:8088/ws/v1/cluster 小技巧如果网关所在节点的HADOOP_CONF_DIR中包含有效的yarn-site.xml则EG_YARN_ENDPOINT可保持默认值NoneYARN 客户端库会自动从配置中定位资源管理器高可用场景同样适用。启动网关并验证部署成果建议将 Jupyter Enterprise Gateway 作为后台服务运行并开启空闲内核回收culling以释放资源详细说明见 launching-eg.mdjupyter enterprisegateway --ip0.0.0.0 --port_retries0 \ --RemoteKernelManager.cull_idle_timeout43200 \ --MappingKernelManager.cull_interval60启动后打开浏览器访问 YARN 资源管理器界面你应该能看到名为 Spark 的 YARN 应用在集群中运行状态为 RUNNING常见问题与避坑指南依赖版本约束Jupyter Enterprise Gateway 要求jupyter_client 7、jupyter_server 2.0、pyzmq 25务必安装在独立的 Python 环境中不要与 Jupyter Notebook/Lab 混装。内核包必须全节点安装Python 和 R 内核包需要安装到集群的每个可用节点pip install ipykernel否则提交任务会失败。端口规划建议使用--port_retries0确保单实例启动避免端口冲突。结语通过在 YARN 集群部署 Jupyter Enterprise Gateway你的团队将获得一个安全、多租户、可弹性扩展的企业级数据科学平台——Spark 任务跑在集群里Notebook 体验保持不变资源利用率却提升了一个数量级。还等什么按照本文步骤动手部署吧【免费下载链接】enterprise_gatewayA lightweight, multi-tenant, scalable and secure gateway that enables Jupyter Notebooks to share resources across distributed clusters such as Apache Spark, Kubernetes and others.项目地址: https://gitcode.com/gh_mirrors/en/enterprise_gateway创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考