Apache Spark™ version update
You can change the Apache Spark™ version to any of the versions supported by Managed Service for Apache Spark™.
Updates and fixes within a version are installed automatically during maintenance.
Get a list of available versions
Note
If you use environments to manage versions, the version is indicated in the base environment name.
- In the management console
, select the folder. - Navigate
to Managed Service for Apache Spark™. - Select a cluster and click
Edit on the top panel. This will open the cluster editing page.
You can see the list of available versions in the Version field.
Before a version upgrade
Make sure the upgrade will not disrupt your applications:
- Check the Apache Spark™ release notes
to learn how upgrades may affect your applications. - Try upgrading the Apache Spark™ version on a test cluster.
Upgrading the version
Apache Spark™ version upgrade method depends on how dependencies are configured in your cluster:
-
If the Packages manually value is selected in the Software configuration method field, change the version in the cluster settings.
Warning
This method is deprecated and will soon become unavailable. We recommend using environments to work with different Apache Spark™ versions.
-
If the Environment value is selected in the Software configuration method field, select the base environment (Apache Spark™ and Python version) or connect a custom environment with your required version to the cluster. The Apache Spark™ version is fixed within the environment and cannot be set separately in the cluster settings.
Cluster uses manually added packages
- In the management console
, select the folder. - Navigate
to Managed Service for Apache Spark™. - Select a cluster and click
Edit on the top panel. - Under Advanced settings, go to the Software configuration method → Packages manually and select an Apache Spark™ version.
- Click Save changes.
If you do not have the Yandex Cloud CLI yet, install and initialize it.
The folder used by default is the one specified when creating the CLI profile. To change the default folder, use the yc config set folder-id <folder_ID> command. You can also specify a different folder for any command using --folder-name or --folder-id.
If you access a resource by its name, the search will be limited to the default folder. If you access a resource by its ID, the search will be global, i.e., through all folders based on access permissions.
To change the Apache Spark™ version:
-
View the description of the CLI command for updating a cluster:
yc managed-spark cluster update --help -
Change the version by running this command:
yc managed-spark cluster update <cluster_name_or_ID> \ --spark-version <Apache_Spark_version>You can get the cluster name and ID with the list of clusters in the folder.
-
Open the current Apache Spark™ configuration file with the infrastructure plan.
Learn how to create this file in Creating a cluster.
-
Edit the
spark_versionparameter in the cluster's description:resource "yandex_spark_cluster" "<cluster_name>" { ... config = { ... spark_version = "<Apache_Spark_version>" ... } ... } -
Make sure the settings are correct.
-
In the command line, navigate to the directory that contains the current Terraform configuration files defining the infrastructure.
-
Run this command:
terraform validateTerraform will show any errors found in your configuration files.
-
-
Confirm updating the resources.
-
Run this command to view the planned changes:
terraform planIf you described the configuration correctly, the terminal will display a list of the resources to update and their parameters. This is a verification step that does not apply changes to your resources.
-
If everything looks correct, apply the changes:
-
Run this command:
terraform apply -
Confirm updating the resources.
-
Wait for the operation to complete.
-
-
-
Get an IAM token for API authentication and put it into an environment variable:
export IAM_TOKEN="<IAM_token>" -
Clone the cloudapi
repository:cd ~/ && git clone --depth=1 https://github.com/yandex-cloud/cloudapiBelow, we assume that the repository contents reside in the
~/cloudapi/directory. -
Create a file named
body.jsonand paste the following code into it:{ "cluster_id": "<cluster_ID>", "update_mask": { "paths": [ "config_spec.spark_version" ] }, "config_spec": { "spark_version": "<Apache_Spark_version>" } }Where:
-
cluster_id: Cluster ID.You can get the cluster ID with the list of clusters in the folder.
-
update_mask: List of parameters to update as an array of strings (paths[]).Format for listing settings
"update_mask": { "paths": [ "<setting_1>", "<setting_2>", ... "<setting_N>" ] }Warning
When you update a cluster, all parameters of the object you are modifying will be reset to their defaults unless explicitly provided in the request. To avoid this, list the settings you want to change in the
update_maskparameter. -
spark_version: Apache Spark™ version.
-
-
Call the ClusterService.Update method, e.g., via the following gRPCurl
request:grpcurl \ -format json \ -import-path ~/cloudapi/ \ -import-path ~/cloudapi/third_party/googleapis/ \ -proto ~/cloudapi/yandex/cloud/spark/v1/cluster_service.proto \ -rpc-header "Authorization: Bearer $IAM_TOKEN" \ -d @ \ spark.api.cloud.yandex.net:443 \ yandex.cloud.spark.v1.ClusterService.Update \ < body.json -
Check the server response to make sure your request was successful.
Cluster uses an environment
Your cluster can use either the base or a custom environment.
The base environment contains only the Apache Spark™ version and Python version. You can select it when creating or updating your cluster.
A custom environment contains the base environment and a list of required pip and deb packages enabling you to install additional libraries and applications in the cluster. You need to specify the list of packages manually. You cannot modify the list of packages and Apache Spark™ version in an existing custom environment. To update the Apache Spark™ version, set up a custom environment with the version you need and connect it to the cluster.
-
Set up an environment with the Apache Spark™ you need using one of these methods:
-
In the management console
, select the folder containing the cluster. -
Navigate
to Managed Service for Apache Spark™. -
Select the cluster and click Edit on the top panel.
-
Under Advanced settings, in the Software configuration method → Environment field, select the environment you set up.
-
Click Save changes.
-
Set up an environment with the Apache Spark™ version you need.
-
Get an IAM token for API authentication and put it into an environment variable:
export IAM_TOKEN="<IAM_token>" -
Clone the cloudapi
repository:cd ~/ && git clone --depth=1 https://github.com/yandex-cloud/cloudapiBelow, we assume that the repository contents reside in the
~/cloudapi/directory. -
Create a file named
body.jsonand paste the following code into it:{ "cluster_id": "<cluster_ID>", "update_mask": { "paths": [ "config_spec.environment_id" ] }, "config_spec": { "environment_id": "<environment_ID>" } }Where:
cluster_id: Cluster ID. You can get it with the list of clusters in the folder.environment_id: ID of the environment with the Apache Spark™ version you need. Call the EnvironmentService/List method to get the custom environment ID or EnvironmentService/ListBase to get the base environment ID.
-
Call the ClusterService/Update method, e.g., via the following gRPCurl
request:grpcurl \ -format json \ -import-path ~/cloudapi/ \ -import-path ~/cloudapi/third_party/googleapis/ \ -proto ~/cloudapi/yandex/cloud/spark/v1/cluster_service.proto \ -rpc-header "Authorization: Bearer $IAM_TOKEN" \ -d @ \ spark.api.cloud.yandex.net:443 \ yandex.cloud.spark.v1.ClusterService.Update \ < body.json -
Check the server response to make sure your request was successful.