General questions about DataSphere
Can I get logs of my operations in Yandex Cloud?
Yes, you can request information about operations with your resources from Yandex Cloud logs. Do it by contacting support
How do I find out what personal data is stored in Yandex Cloud?
To find out what personal data is stored in Yandex Cloud, contact technical support. You can also request to have your personal data completely deleted
If a project fails to open, check your current quota usage
What do I do if I cannot install a package in my project or do not have internet access?
Internet connectivity issues may occur if you enabled a subnet that is not configured for Internet access in your project. For example, you may get a ConnectTimeoutError error while installing a package with pip:
Defaulting to user installation because normal site-packages is not writeable
WARNING: Retrying (Retry(total=4, connect=None, read=None, redirect=None, status=None)) after connection broken by 'ConnectTimeoutError(<HTTPSConnection(host='pypi.org', port=443) at 0x7f42c810ca00>, 'Connection to pypi.org timed out. (connect timeout=15)')': /simple/langchain/
ERROR: Could not find a version that satisfies the requirement langchain (from versions: none)
ERROR: No matching distribution found for langchain
If you need a subnet for your project, set up an NAT gateway to enable internet access for this subnet.
You can change or disable the subnet in the project settings.
Can I close a tab with a notebook?
Yes, you can. If you close the notebook tab, current executions will continue, all variables and computation results will be saved, but the output will not be saved for the executions that finished while the notebook was closed.
After completion of all the running computations, the VM will be assigned to the notebook for three hours. You can change this value in the project settings.
How do I specify the configuration type for my project?
You can select a computing resource configuration when you first run computations in the DataSphere notebook. The minimum available configuration is c1.4 (4 vCPUs).
If I delete a running cell, will computations stop?
No, they will not. Computations will continue even if you delete a cell from the notebook. Before deleting a cell, make sure to stop it. If you have deleted a running cell, stop running calculations. To do this, select File ⟶ Stop IDE executions in JupyterLab or click Stop JupyterLab and VM in the Running executions widget on the project page.
How do I clear a cell's outputs?
Select Edit ⟶ Clear All Outputs in JupyterLab or right-click on any cell and select Clear All Outputs. If you choose the second option, the outputs will only be reset for the current session.
Does DataSphere support scheduled cell runs?
You can run scheduled calculations by rerunning the DataSphere Jobs jobs and integrating them with Yandex Managed Service for Apache Airflow™.
You can also use Yandex Cloud Functions to automatically initiate notebook execution using the DataSphere API. For a detailed description of regular runs, see this guide.
My browser cannot open a DataSphere project in the IDE. How can I fix this?
The project is still loading. If the loading icon persists, wait up to ten minutes. Loading a large project with many files can take this long. You can track your browser network activity on the Network tab in your browser's developer tools. To open the developer tools, press F12.
The browser blocks access to the IDE. In this case, you may see a blank page or a Web page temporarily unavailable, Error 503, or Deadline Exceeded message. When opening a project in an IDE, DataSphere redirects your request to its own host with JupyterLab. Modern browsers may block this redirect if you use enhanced privacy settings, including incognito mode. To open a project in an IDE, disable the blocking settings:
- Chrome: Allow using third-party cookies.
- Safari: Disable Website tracking: Prevent cross-site tracking under Preferences → Privacy.
- Yandex Browser: Allow using third-party cookies for DataSphere in the browser settings under Sites → Advanced site settings.
- Firefox: Click the shield icon in the address bar and disable Enhanced Tracking Protection.
If your project fails to open, make sure none of the quotas has reached its limit. Navigate to the Communities tab, select a community, and open the Quotas tab.
My browser asks me to grant access to a JupyterLab host. How do I grant it?
The message is triggered by an experimental option in Google Chrome, which implements the storage access API. To disable it, type chrome://flags in the browser address bar, find Storage Access API in the search bar below, and change the option status to Disabled.
How do I deploy a Hugging Face model in DataSphere?
Some libraries download models to predefined folders by default. Models may not be available for import if the folder they were downloaded to is not located in the project repository. To avoid this, choose a correct download directory and specify it when importing a model:
cache_dir="/home/jupyter/datasphere/project/huggingface_cache_dir/"
config = AutoConfig.from_pretrained("<model_name>", cache_dir=cache_dir)
model = AutoModel.from_pretrained("<model_name>", config=config, cache_dir=cache_dir)
To avoid specifying the directory's path every time, you can provide it in an environment variable. Make sure you do this at the very start of the notebook, prior to importing libraries:
import os
os.environ['TRANSFORMERS_CACHE'] = '/home/jupyter/datasphere/project/huggingface_cache_dir/'
In addition, you can configure a model to operate in offline mode by referring to the official Hugging Face documentation
Why do I get an Access Denied: Spec gX.X is not available for your cloud error?
This error means the selected GPU configuration is not available in your cloud. Configurations marked with footnote 1 become available after switching to paid usage and topping up the billing account. The required top-up amount is specified in this article.
After topping up your balance:
- Stop computations through File ⟶ Stop IDE executions and exit the project.
- Reopen the project in the management console
.
To learn how resources are charged within projects, see DataSphere pricing policy.
Why does code in project cells take a long time to run?
When running code, a Preparing <configuration_name> instance message may be displayed for a long time. The possible causes may include the following:
- The notebook has many cells, or its cells contain a large amount of code. Fully loading a project that contains hundreds of cells may take more than ten minutes. If feasible, try to reduce the number of cells in your project or opt for a higher-performance computing configuration.
- When running code for the first time, DataSphere provisions and starts a virtual machine with the selected resource configuration. This may take a while. To reduce startup time, use a higher-performance resource configuration.
Why do I get a Servant not allocated error when running code?
When running code, you may see the following messages:
Preparing <configuration_name> instance
Execute error: Servant <configuration_name> not allocated: Internal Error
This error may occur when available computing resources are temporarily insufficient. Try using a different configuration or rerun the code later.
If your project uses a custom Docker image, check whether the issue persists on a public image.
What should I do if I get a Device or resource busy error when installing a library?
When installing a library, you may get this error:
ERROR: Could not install packages due to an OSError: [Errno 16] Device or resource busy: '.nfs0000000000009d3a00000018'
This is a system message related to resources. To fix this error:
- Restart the kernel by selecting Kernel ⟶ Restart Kernel.
- Stop computations by selecting File ⟶ Stop IDE executions in JupyterLab or clicking Stop JupyterLab and VM in the Running executions widget on the project page.