Repository navigation
Conversation
|
Example of Details
Files produced by a run on my workstation using overlayfs and the dryrun mode Details
|
|
|
||
| source "${venv_dir}/bin/activate" || error "Failed to activate the Python virtual environment in ${venv_dir}!" | ||
| pip install --upgrade pip || error "Failed to upgrade pip in the Python virtual environment in ${venv_dir}!" | ||
| pip install "${EB_JPT_PACKAGE}[cli] @ git+${EB_JPT_KERNELS_REPO}@${EB_JPT_KERNELS_COMMIT}" |
There was a problem hiding this comment.
The script is fine, but this alone makes me very concerned that we are opening a security hole in the ingestion process. We talked a bit about this and it may be better to move a variant of this script to https://github.com/EESSI/software-layer-scripts and have a way to trigger that script there.
The trigger would ideally happen in the software-layer, where we could add CI for a push to main which runs the scripts to see if any new/changed kernels would be created (failing CI if they are, or potentially even opening the PR to fix it)
There was a problem hiding this comment.
Will try and prepare a PoC with the other approach.
The alternative would be to move the pip install part out of the cvmfs transaction block but that leaves the concern of a supply-chain attack on the package/deps themselves.
In the meanwhile i think the other points in the PR on where we want to put the kernels and how to deal the host injection should still be discussed.
The accelerator handling might be solved by moving to software layer since there we are running in the proper environment and any accel-specific module should already be picked up.
Also i guess the general overview of how the frontend looks like (eg kernel display names) is worth discussing


Changes
Use the CLI from https://github.com/Crivella/easybuild_jupyter_kernels to generate hardcoded jupyter kernels for every ARCH of the EESSI version being updated by a software tarball ingenstion
NOTES
I am mimicking what the procedure for updating the Lmod caches does to run once on ingestion for every architecture.
There is some trickery required to get the correct module path for the arch we want but still use binaries from the compat layer (eg
uname) to allowmodule loadto work.For now the kernels are stored under
<CVMFS_REPO>/versions/<CURRENT_VERSION>/software/linux/<ARCH_PATH>/.jupyter/kernelswhich means that one would have to add the
<CVMFS_REPO>/versions/<CURRENT_VERSION>/software/linux/<ARCH_PATH>/.jupyterto theirJUPYTER_PATHto pick them up when starting a serverTime for running the new script is similar to the generation of the Lmod caches since it has to run for every arch (unless we want to assume that the paths addede by the various archs are all the same just swapping the arch part of the path)
Other changes
ingest-tarball.shto be ran in a ~dry-runmode where thecvmfs_servercalls are mocked (to test eg locally by mounting an overlayfs on top of EESSI)TODO
EBKernelSpecManagerfrom https://github.com/Crivella/easybuild_jupyter_kernelsmodule loadcommands inside the notebook/console itself. See EG what i've done for the AITW docker extension https://github.com/AI-TranspWood/jupyter-notebooks/blob/main/notebooks/SurrogateStiffness.ipynb