<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Jupyter Blog - HPC</title><link href="https://blog.jupyter.org/" rel="alternate"/><link href="https://blog.jupyter.org/feeds/tag-hpc.atom.xml" rel="self"/><id>https://blog.jupyter.org/</id><updated>2026-04-20T18:04:00+00:00</updated><subtitle>Jupyter for high-performance and parallel computing, for example, at supercomputing centers and scientific facilities.</subtitle><entry><title>Exploring Petabytes of the Night Sky — Jupyter Notebooks at NOIRLab’s Astro Data Lab Science Platform</title><link href="https://blog.jupyter.org/posts/2026/exploring-petabytes-of-the-night-sky-jupyter-notebooks/" rel="alternate"/><published>2026-04-20T18:04:00+00:00</published><updated>2026-04-20T18:04:00+00:00</updated><author><name>Robert Nikutta</name></author><id>tag:blog.jupyter.org,2026-04-20:/posts/2026/exploring-petabytes-of-the-night-sky-jupyter-notebooks/</id><summary type="html">&lt;p&gt;Imagine querying 420+ billion rows of astronomical catalog data — spanning 30 major sky surveys, observed over decades with telescopes on three continents — from a Jupyter notebook in your browser in seconds. No download. No HPC allocation request. No waiting.&lt;/p&gt;</summary><content type="html">&lt;p&gt;Imagine querying 420+ billion rows of astronomical catalog data — spanning 30 major sky surveys, observed over decades with telescopes on three continents — from a Jupyter notebook in your browser in seconds. No download. No HPC allocation request. No waiting.&lt;/p&gt;
&lt;p&gt;That is what 4,800+ astronomers in over 90 countries can do every day at the &lt;a href="https://datalab.noirlab.edu"&gt;Astro Data Lab&lt;/a&gt; science platform. Data Lab is operated by &lt;a href="https://noirlab.edu"&gt;NSF NOIRLab&lt;/a&gt;, the National Optical-Infrared Astronomy Research Laboratory, headquartered in Tucson, Arizona, with observatories in Arizona, Hawai’i, and Chile. Since its public launch in June 2017, Astro Data Lab has quietly become one of the largest deployments of Jupyter notebooks in professional science — and a case study in what happens when you bring the compute to the data instead of the other way around.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="World map of 2024 number of data queries at Astro Data Lab by country (log-scale color): 64.2 million queries from 72 countries that year." src="https://blog.jupyter.org/posts/2026/exploring-petabytes-of-the-night-sky-jupyter-notebooks/images/001-1_bllZBblDcMeLN2fxIS9PsA.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;World map of 2024 number of data queries at Astro Data Lab by country (log-scale color): 64.2 million queries from 72 countries that year.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="the-data-problem-astronomy-had-to-solve"&gt;The Data Problem Astronomy Had to Solve&lt;/h2&gt;
&lt;p&gt;Modern sky surveys are data machines. The &lt;a href="https://www.darkenergysurvey.org/"&gt;Dark Energy Survey&lt;/a&gt; cataloged 690 million objects. &lt;a href="https://www.esa.int/Science_Exploration/Space_Science/Gaia_overview"&gt;Gaia&lt;/a&gt; measured positions and motions for 1.8 billion stars. The &lt;a href="https://www.legacysurvey.org"&gt;DESI Legacy Surveys&lt;/a&gt; cover 20,000 square degrees, nearly half of the full sky, in three optical bands. And the upcoming Rubin Observatory’s &lt;a href="https://rubinobservatory.org/explore/how-rubin-works/lsst"&gt;Legacy Survey of Space and Time&lt;/a&gt; (LSST) will generate roughly 10 million transient alerts &lt;em&gt;per night&lt;/em&gt; starting later this year.&lt;/p&gt;
&lt;p&gt;Traditional astronomy workflows begin with downloading relevant data to a local computer and to use locally installed specialized software tools to process and analyze the data. However, downloading these catalogs to a local machine is now often physically impossible. A single survey’s measurements table can exceed the combined disk space of an entire research group. And even if you could download it, the computing resources needed to query it efficiently at scale requires infrastructure most astronomers don’t have.&lt;/p&gt;
&lt;p&gt;The answer the community converged on, like many industries dealing with big data: bring the compute to the data. Host the catalogs in databases, co-locate a computing environment next door, and give scientists a familiar interface to work in. That interface, increasingly, is a Jupyter notebook.&lt;/p&gt;
&lt;h2 id="astro-data-lab-jupyter-at-the-observatory"&gt;Astro Data Lab: Jupyter at the Observatory&lt;/h2&gt;
&lt;p&gt;Astro Data Lab was conceived in 2014 and went public in June 2017, originally built to support data releases from the Dark Energy Survey — a few terabytes of catalogs and tens of terabytes of images. We imagined a ceiling of roughly 500 users. We were wrong, in the best way.&lt;/p&gt;
&lt;p&gt;Today the platform hosts:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;About 420 billion catalog rows&lt;/strong&gt; across 30+ major astronomical surveys (DES, Legacy Surveys, DESI, NOIRLab Source Catalog, SDSS, Gaia, unWISE, SMASH, S-PLUS, VHS, 2MASS, and dozens more)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;31 million spectra&lt;/strong&gt; via SPARCL, our spectral access service (DESI DR1+EDR, SDSS/BOSS DR17)&lt;/li&gt;
&lt;li&gt;Petabytes of images, accessible through a &lt;a href="https://www.ivoa.net/documents/SIA/"&gt;Simple Image Access&lt;/a&gt; service and cutout API&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Over 4,800 registered users&lt;/strong&gt; from over 90 countries, who submit tens of millions of data queries each year&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Every registered user gets a persistent JupyterHub environment with the full astronomy Python stack pre-loaded — Astropy, NumPy, SciPy, Matplotlib, Pandas, Scikit-learn — and our own astro-datalab client library. The library provides core services, for instance auth and DB queries:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;dl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;authClient&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;queryClient&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;getpass&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;getpass&lt;/span&gt;

&lt;span class="c1"&gt;# Log in&lt;/span&gt;
&lt;span class="n"&gt;token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;authClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;login&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;username: &amp;quot;&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;&lt;span class="n"&gt;getpass&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;password: &amp;quot;&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="c1"&gt;# Query 10 objects from the NOIRLab Source Catalog near a sky position&lt;/span&gt;
&lt;span class="c1"&gt;# Right Ascension (RA) = 150.12 degrees&lt;/span&gt;
&lt;span class="c1"&gt;# Declination (Dec) = 2.21 degrees&lt;/span&gt;
&lt;span class="c1"&gt;# Search radius = 0.05 degrees&lt;/span&gt;
&lt;span class="c1"&gt;# q3c (Quad Tree Cube) is a spatial indexing scheme for Postgres&lt;/span&gt;
&lt;span class="c1"&gt;# gmag and rmag are the g-band and r-band magnitudes of objects&lt;/span&gt;
&lt;span class="c1"&gt;# in the NOIRLab Source Catalog Data Release 2, ‘object’ table.&lt;/span&gt;
&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;queryClient&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="sd"&gt;&amp;quot;&amp;quot;&amp;quot;SELECT ra, dec, gmag, rmag FROM nsc_dr2.object&lt;/span&gt;
&lt;span class="sd"&gt;       WHERE q3c_radial_query(ra, dec, 150.12, 2.21, 0.05) LIMIT 10&amp;quot;&amp;quot;&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;fmt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;pandas&amp;quot;&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;It’s as simple as that. The query runs on the database server next to the data; only the result set crosses the network.&lt;/p&gt;
&lt;h2 id="use-case-1-seeing-the-sky-inside-a-notebook-with-aladinlite"&gt;Use Case 1 — Seeing the Sky Inside a Notebook with AladinLite&lt;/h2&gt;
&lt;p&gt;One of the most immediate joys of working with astronomical data is visualization: not just numbers in a table, but &lt;em&gt;where things are in the sky&lt;/em&gt;, what the images look like, and how your query results relate to the underlying survey footprint.&lt;/p&gt;
&lt;p&gt;We’ve integrated &lt;a href="https://aladin.cds.unistra.fr/AladinLite/"&gt;AladinLite v3&lt;/a&gt; — the interactive sky atlas from Centre de Données Astronomiques de Strasbourg (CDS) — directly into the notebook environment via the &lt;a href="https://github.com/cds-astro/ipyaladin"&gt;ipyaladin&lt;/a&gt; widget. With a handful of lines, astronomers can embed a fully interactive sky viewer in a notebook cell or next to their notebook in a “sidecar”, and overlay their own data on top of real survey imagery:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;astropy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;units&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;astropy.table&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Table&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;astropy.coordinates&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SkyCoord&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;ipyaladin&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Aladin&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;sidecar&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Sidecar&lt;/span&gt;

&lt;span class="c1"&gt;# Instantiate the Aladin interactive sky viewer &lt;/span&gt;
&lt;span class="n"&gt;aladin&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Aladin&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;full_screen&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="kc"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;Sidecar&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;title&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;aladin_output&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="n"&gt;anchor&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;split-right&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;display&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;aladin&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# globular cluster NGC 1851 (RA, Dec)&lt;/span&gt;
&lt;span class="n"&gt;aladin&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;SkyCoord&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;78.52809&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deg&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mf"&gt;40.04656&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;aladin&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;coo_frame&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;quot;ICRSd&amp;quot;&lt;/span&gt;  &lt;span class="c1"&gt;# set coordinate frame to ICRS, angles in deg&lt;/span&gt;
&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# race condition&lt;/span&gt;
&lt;span class="n"&gt;aladin&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;fov&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.4&lt;/span&gt;  &lt;span class="c1"&gt;# set field of view to 0.4 degrees&lt;/span&gt;

&lt;span class="c1"&gt;# Overlay catalog query results as circle markers&lt;/span&gt;
&lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;Table&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;from_pandas&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;df&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# e.g., from a previous query around NGC 1851&lt;/span&gt;
&lt;span class="n"&gt;aladin&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;add_table&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;circle&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="n"&gt;source_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="n"&gt;color&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;green&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The result is a pannable, zoomable sky viewer — right in the notebook — with your query results overlaid as green circles on the actual sky image of a globular cluster (see figure below). Users can overlay MOCs (Multi-Order Coverage maps, which encode survey footprints), user-generated catalogs from a prior query, or any Virtual Observatory-standard data source.&lt;/p&gt;
&lt;p&gt;This capability turns what was once a static plot into an exploratory tool: zoom into a cluster, click on a source, cross-match on the fly. For students and scientists unfamiliar with a dataset, it is often the fastest path from “I have a list of objects” to “I understand where they are and what I’m looking at.”&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/astro-datalab/notebooks-latest/blob/master/04_HowTos/Aladin/ipyaladin_MOC.ipynb"&gt;AladinLite integration is now active in our notebook library&lt;/a&gt;, with full deployment into the new Data Lab Web Portal on the roadmap for later this year.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="AladinLite v3 sky viewer inside a Jupyter notebook, showing a globular cluster with catalog query results in the outskirts of the cluster overlaid as green circles." src="https://blog.jupyter.org/posts/2026/exploring-petabytes-of-the-night-sky-jupyter-notebooks/images/002-1_3kJd4UCGCp0IpTrcL7fxqA.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;&lt;a href="https://github.com/astro-datalab/notebooks-latest/blob/master/04_HowTos/Aladin/ipyaladin_globular_cluster.ipynb"&gt;AladinLite v3 sky viewer inside a Jupyter notebook&lt;/a&gt;, showing a globular cluster with catalog query results in the outskirts of the cluster overlaid as green circles.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="use-case-2-stacking-galaxy-spectra-with-sparcl"&gt;Use Case 2 — Stacking Galaxy Spectra with SPARCL&lt;/h2&gt;
&lt;p&gt;Spectroscopy — measuring how much light a star or a galaxy emits at each wavelength — is one of astronomy’s most powerful tools. But individual spectra are often noisy. The signal-to-noise ratio of a single optical spectrum for a faint galaxy can be too low to measure the emission lines that encode star formation rate, gas chemical content (Oxygen, Nitrogen, etc.) or the even more subtle absorption lines that create small wiggles in the shape of the spectrum, yet encapsulate crucial information such as the mass and age of the stars making up a galaxy.&lt;/p&gt;
&lt;p&gt;One trick that astronomers have used for decades: combining or “stacking” spectra. Average hundreds of spectra together, and the noise level reduces while the signal builds up. What was invisible in a single spectrum becomes unmistakable in the stack. While the concept is simple, reading and manipulating large numbers of spectra can be time consuming or cumbersome.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://astrosparcl.datalab.noirlab.edu"&gt;SPARCL&lt;/a&gt; (SPectra Analysis and Retrievable Catalog Lab) makes this possible at scale directly in a notebook. With &lt;code&gt;sparclclient&lt;/code&gt;, users can currently search 31 million spectra by redshift range, target type, and survey, then retrieve flux arrays and wavelength grids ready for stacking:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;sparcl.client&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SparclClient&lt;/span&gt;

&lt;span class="c1"&gt;# Instantiate the SPARCL client (connected to production server)&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;SparclClient&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="c1"&gt;# Find SDSS spectra of galaxies in a redshift slice 0.1&amp;lt;z&amp;lt;0.3&lt;/span&gt;
&lt;span class="n"&gt;found&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;find&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;outfields&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;sparcl_id&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;ra&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;dec&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;redshift&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;spectype&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;constraints&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;spectype&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;GALAXY&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                 &lt;span class="s1"&gt;&amp;#39;redshift&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mf"&gt;0.1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
                 &lt;span class="s1"&gt;&amp;#39;data_release&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;SDSS-DR17&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;]},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Retrieve flux, wavelength, and inverse-variance arrays&lt;/span&gt;
&lt;span class="n"&gt;retrieved&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;retrieve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;found&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ids&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;include&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;flux&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;wavelength&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;ivar&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;In our &lt;a href="https://github.com/astro-datalab/notebooks-latest/blob/master/03_ScienceExamples/SpectralStacking/SpectralStacking_SDSS.ipynb"&gt;SpectralStacking_SDSS&lt;/a&gt; science example notebook, users first stack a small number of galaxy spectra (N=5) in eight bins of astrophysical color &lt;em&gt;g&lt;/em&gt;−&lt;em&gt;r&lt;/em&gt; (green and red filters), revealing trends from blue spectra with emission lines to red spectra with absorption lines but with noisy spectra. Then users stack hundreds of galaxy spectra for the same bins of color &lt;em&gt;g&lt;/em&gt;−&lt;em&gt;r&lt;/em&gt; and obtain much cleaner spectra where the small wiggles are now real astrophysical features and no longer buried in the noise.&lt;/p&gt;
&lt;p&gt;The spectral rainbows below — N=5 then N=200 stacked galaxy spectra color-coded in bins of astrophysical color &lt;em&gt;g&lt;/em&gt;−&lt;em&gt;r&lt;/em&gt; — are each a single output cell from this notebook, generated entirely within the Data Lab environment.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="SPARCL spectral rainbows: top panel shows N=5 galaxy spectra stacked and color-coded by g−r color, spanning wavelengths 3750–6450 Ångstrom. Bottom panel shows the same exercise but with N=200 galaxy spectra per bin, greatly enhancing the signal-to-noise ratio." src="https://blog.jupyter.org/posts/2026/exploring-petabytes-of-the-night-sky-jupyter-notebooks/images/003-1_NZpW8VO92rpjjqyR47BywQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;SPARCL spectral rainbows: top panel shows N=5 galaxy spectra stacked and color-coded by g−r color, spanning wavelengths 3750–6450 Ångstrom. Bottom panel shows the same exercise but with N=200 galaxy spectra per bin, greatly enhancing the signal-to-noise ratio.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="use-case-3-variable-stars-and-the-coming-flood-of-time-domain-data"&gt;Use Case 3 — Variable Stars and the Coming Flood of Time-Domain Data&lt;/h2&gt;
&lt;p&gt;Not all astronomical data is a static snapshot of the sky. Many of the most scientifically rich phenomena — pulsating stars, transiting exoplanets, exploding supernovae, gravitational lensing events — reveal themselves through &lt;em&gt;change&lt;/em&gt; over time.&lt;/p&gt;
&lt;p&gt;Among the most useful calibration tools in astrophysics are RR Lyrae stars: old, low-mass stars that pulsate with periods of 0.2–1 day and a brightness variation that traces their distance. Finding and characterizing them across millions of square degrees of sky requires querying multi-epoch photometry catalogs, computing period statistics, and folding light curves — all tasks that fit naturally in a notebook workflow.&lt;/p&gt;
&lt;p&gt;Our &lt;a href="https://github.com/astro-datalab/notebooks-latest/blob/master/03_ScienceExamples/TimeSeriesAnalysisRrLyraeStar/TimeSeriesAnalysisOfRrLyraeStar.ipynb"&gt;TimeSeriesAnalysisRrLyraeStar&lt;/a&gt; notebook demonstrates the full pipeline: query the SMASH DR2 catalog for stars with high photometric variability, run a Lomb-Scargle periodogram on the light curve, identify the dominant period, and phase-fold the observations to reveal the characteristic sawtooth pulsation profile:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;astropy.timeseries&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;LombScargle&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;numpy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;as&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;np&lt;/span&gt;

&lt;span class="n"&gt;ls&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;LombScargle&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;y&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# time and magnitude from a previous query&lt;/span&gt;
&lt;span class="n"&gt;frequency&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;power&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ls&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;autopower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;period&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;1.&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;frequency&lt;/span&gt; &lt;span class="c1"&gt;# period is the inverse of frequency&lt;/span&gt;
&lt;span class="n"&gt;best_period&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;period&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;argmax&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;power&lt;/span&gt;&lt;span class="p"&gt;)]&lt;/span&gt;
&lt;span class="n"&gt;phase&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;t&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;best_period&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;  &lt;span class="c1"&gt;# folded timeseries = light curve&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The resulting phase-folded light curve shown in the figure below is clean, precise, and immediately recognizable to any variable-star astronomer — produced entirely from archival survey data without a single new observation.&lt;/p&gt;
&lt;p&gt;This kind of workflow is also a proving ground for the upcoming &lt;a href="https://rubinobservatory.org/"&gt;Vera C. Rubin Observatory&lt;/a&gt;’s &lt;a href="https://rubinobservatory.org/explore/how-rubin-works/lsst"&gt;Legacy Survey of Space and Time&lt;/a&gt; (LSST). When Rubin begins operations and delivers 10 million nightly alerts, the only workflows that will scale are ones already designed to run against large databases or specialized file systems, in shared computing environments, with notebook-native tooling. Astro Data Lab users are building those workflows today.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Phase-folded RR Lyrae light curve from the TimeSeriesAnalysisRrLyraeStar notebook, showing characteristic sawtooth pulsation with a 0.65 day period." src="https://blog.jupyter.org/posts/2026/exploring-petabytes-of-the-night-sky-jupyter-notebooks/images/004-1_JDuoMo_2SgPJ-b_md9nslg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Phase-folded RR Lyrae light curve from the TimeSeriesAnalysisRrLyraeStar notebook, showing characteristic sawtooth pulsation with a 0.65 day period.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="the-notebook-ecosystem"&gt;The Notebook Ecosystem&lt;/h2&gt;
&lt;p&gt;The three use cases above are drawn from our library of &lt;strong&gt;80+ open-source Jupyter notebooks&lt;/strong&gt; at &lt;a href="https://github.com/astro-datalab/notebooks-latest"&gt;github.com/astro-datalab/notebooks-latest&lt;/a&gt;. The library is organized into six sections:&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Directory&lt;/th&gt;
&lt;th&gt;Contents&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;01_GettingStartedWithDataLab/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Authentication, dataset discovery, first queries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;02_DataAccessOverview/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;More advanced queries, image searches, etc.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;03_ScienceExamples/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Many complete science cases (stellar streams, dwarf galaxies, large-scale structure, SED fitting, …)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;04_HowTos/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Service-specific tutorials (SPARCL, SIA image cutouts, cross-matching, file storage)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;05_Contrib/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Community-contributed notebooks (ANTARES alert broker, user science cases, etc.)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;06_EPO/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Education &amp;amp; public outreach (Teen Astronomy Cafe, La Serena School for Data Science)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;All notebooks are open-source and community contributions are welcome via pull request. We use them as living teaching materials in workshops at Astronomical Data Analysis Software &amp;amp; Systems (ADASS) and American Astronomical Society (AAS) conferences, summer schools, and university courses around the world. We have also recently translated most of our notebooks to the &lt;a href="https://github.com/astro-datalab/notebooks-latest-es"&gt;Spanish language&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;In the coming months we will launch a &lt;strong&gt;tagged, searchable notebook gallery&lt;/strong&gt; — filterable by science topic, &lt;a href="https://astrothesaurus.org/"&gt;Unified Astronomy Thesaurus&lt;/a&gt; (UAT) keywords, target audience, and difficulty level. The pilot framework was developed by two summer students working with the team.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Footprints of 24 sky survey datasets hosted at Astro Data Lab. This montage shows the wide variety of astronomical surveys, with some that cover the full sky, others focusing on the Milky Way (central plane), and yet others studying the extragalactic regions beyond the Milky Way. We ensure that each survey is represented in at least one of our example notebooks." src="https://blog.jupyter.org/posts/2026/exploring-petabytes-of-the-night-sky-jupyter-notebooks/images/005-1_SvI9n6mkoQ5th35aClATWQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Footprints of 24 sky survey datasets hosted at Astro Data Lab. This montage shows the wide variety of astronomical surveys, with some that cover the full sky, others focusing on the Milky Way (central plane), and yet others studying the extragalactic regions beyond the Milky Way. We ensure that each survey is represented in at least one of our &lt;a href="https://github.com/astro-datalab/notebooks-latest/"&gt;&lt;em&gt;example notebooks&lt;/em&gt;&lt;/a&gt;.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="looking-ahead"&gt;Looking Ahead&lt;/h2&gt;
&lt;p&gt;Nine years in, the Astro Data Lab science platform is evolving on several fronts simultaneously.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GPU computing.&lt;/strong&gt; We are deploying a GPU node, which will be connected to the Jupyter notebook service. This opens deep learning and large-scale ML workflows in the same notebook environment where the data lives.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;An AI assistant.&lt;/strong&gt; Our first-ever user survey, conducted in September 2025, ranked an in-notebook AI assistant as one of the top requested features. We are actively exploring what responsible, science-aware AI assistance looks like in this context — helping users construct SQL/ADQL queries, navigate datasets, and debug notebook code, without hallucinating catalog column names. jupyter-ai might come in very handy here.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;New integrated Web Portal.&lt;/strong&gt; Our Data Explorer — an integrated web interface combining catalog browsing, query execution, image cutouts, spectral search, and job status monitoring — was rolled out last year. Some of the next milestones include integration of AladinLite into the portal, bringing the sky-visualization capability described above out of the notebook and into the browser-native interface, and a new integrated positional cross-matching service.&lt;/p&gt;
&lt;h2 id="try-it"&gt;Try It&lt;/h2&gt;
&lt;p&gt;The full notebook library is open-source: &lt;a href="https://github.com/astro-datalab/notebooks-latest/"&gt;github.com/astro-datalab/notebooks-latest&lt;/a&gt;. Community notebook contributions are welcome — see &lt;a href="https://github.com/astro-datalab/notebooks-latest/blob/master/CONTRIBUTING"&gt;CONTRIBUTING&lt;/a&gt; in the repository. You can also run all notebooks locally, after installing the Data Lab command-line client and Python module: &lt;code&gt;pip install astro-datalab&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;Astro Data Lab also offers a JupyterLab environment as a service to the broad astronomy community — students, researchers, educators, and citizen scientists. &lt;a href="https://datalab.noirlab.edu/account/register/"&gt;Registration&lt;/a&gt; takes just a moment at &lt;a href="https://datalab.noirlab.edu"&gt;datalab.noirlab.edu&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Questions and feedback: &lt;a href="mailto:datalab@noirlab.edu"&gt;datalab@noirlab.edu&lt;/a&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="mailto:robert.nikutta@noirlab.edu"&gt;Robert Nikutta&lt;/a&gt; is a scientist at NSF NOIRLab’s Community Science and Data Center, and lead of the Astro Data Lab science platform. &lt;a href="mailto:stephanie.juneau@noirlab.edu"&gt;Stéphanie Juneau&lt;/a&gt; is an associate astronomer at CSDC and lead of the SPARCL spectroscopy initiative. The platform is the work of the full &lt;a href="https://datalab.noirlab.edu/about/people"&gt;Astro Data Lab team&lt;/a&gt;, past and present.&lt;/em&gt;&lt;/p&gt;
</content><category term="HPC"/><category term="science"/></entry><entry><title>IPython Parallel in 2021</title><link href="https://blog.jupyter.org/posts/2021/ipython-parallel-in-2021/" rel="alternate"/><published>2021-11-30T07:55:00+00:00</published><updated>2021-11-30T07:55:00+00:00</updated><author><name>Min RK</name></author><id>tag:blog.jupyter.org,2021-11-30:/posts/2021/ipython-parallel-in-2021/</id><summary type="html">&lt;p&gt;Updates on IPython Parallel; new features, future direction&lt;/p&gt;</summary><content type="html">&lt;p&gt;&lt;em&gt;Updates on IPython Parallel; new features, future direction&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;This post describes work funded by Bodo, Inc.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://ipyparallel.readthedocs.io/"&gt;IPython Parallel&lt;/a&gt;’s a bit of an odd duck in the parallel computing space. In a world with &lt;a href="https://dask.org/"&gt;dask&lt;/a&gt;, &lt;a href="https://ray.io/"&gt;ray&lt;/a&gt;, &lt;a href="https://docs.bodo.ai/latest/source/getting_started.html"&gt;bodo&lt;/a&gt;, &lt;a href="https://spark.apache.org/docs/latest/api/python/index.html"&gt;pyspark&lt;/a&gt;, and other parallel computing tools, what is IPython Parallel’s role in 2021?&lt;/p&gt;
&lt;h2 id="context"&gt;Context&lt;/h2&gt;
&lt;p&gt;IPython Parallel began (in 2006!) as a natural extension of the &lt;a href="https://jupyter-client.readthedocs.io/en/stable/messaging.html#general-message-format"&gt;Jupyter messaging protocol&lt;/a&gt;(&lt;em&gt;though it predates the Jupyter name by a few years&lt;/em&gt;): when you have a protocol for &lt;a href="https://en.wikipedia.org/wiki/Read%E2%80%93eval%E2%80%93print_loop"&gt;REPL&lt;/a&gt;-style remote code execution, what can you do with &lt;em&gt;multiple&lt;/em&gt; remote execution environments? This question led to the development two basic models for parallel execution in IPython Parallel:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;“multiplexed execution” — where you send code explicitly to different workers, and&lt;/li&gt;
&lt;li&gt;“load-balanced execution” — where you tell IPython Parallel what to run, and a scheduler takes care of assigning each task to an available worker.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="what-ipython-parallel-doesnt-do"&gt;What IPython Parallel doesn’t do&lt;/h2&gt;
&lt;p&gt;To IPython, your tasks are black boxes (either Python functions or blocks of Python code as text) and you are in complete control of where and when your tasks run, as well as any dependencies or side effects they may have.&lt;/p&gt;
&lt;p&gt;IPython Parallel specifically doesn’t and won’t do several things that other tools might:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;infer dependencies or relationships between tasks (execution graph)&lt;/li&gt;
&lt;li&gt;manage data distribution or locality&lt;/li&gt;
&lt;li&gt;construct parallel algorithms or represent ‘natively’ parallel distributed data structures, such as distributed arrays or data frames&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Everything is very explicit and implemented at the client-level in IPython Parallel, meaning that while &lt;em&gt;it&lt;/em&gt; doesn’t do these things, &lt;em&gt;you&lt;/em&gt; can. It also means that many of the things that have overhead costs associated with parallel computing (communication, data movement) have particularly high costs in IPython Parallel.&lt;/p&gt;
&lt;p&gt;IPython Parallel doesn’t hide anything from you, for better &lt;em&gt;and&lt;/em&gt; worse.&lt;/p&gt;
&lt;h2 id="what-ipython-parallel-does"&gt;What IPython Parallel does&lt;/h2&gt;
&lt;p&gt;IPython Parallel &lt;strong&gt;makes &lt;em&gt;explicit&lt;/em&gt; parallel computations interactive&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;IPython Parallel is not just parallel Python, it’s parallel &lt;em&gt;IPython&lt;/em&gt;. That means you have all the power of IPython’s inspection, interactivity, magics, and debugging available on all of your distributed workers. This makes it especially well-suited to prototyping and experimentation.&lt;/p&gt;
&lt;p&gt;IPython Parallel also presents standard APIs such as &lt;a href="https://ipyparallel.readthedocs.io/en/8.0.0/examples/Futures.html#Executors"&gt;Python Executors&lt;/a&gt;, compatible with many other implementations, to make it easy to migrate to &lt;em&gt;and from&lt;/em&gt; IPython Parallel, enabling developers to write code that uses a single multicore laptop or a thousand cores on an HPC cluster or cloud.&lt;/p&gt;
&lt;h2 id="ipython-parallel-in-2021"&gt;IPython Parallel in 2021&lt;/h2&gt;
&lt;figure&gt;
&lt;img alt="Interactive progress across parallel engines. Those progress bars are interactive widgets running locally, and on each remote engine!" src="https://blog.jupyter.org/posts/2021/ipython-parallel-in-2021/images/001-1_5W4lCfV_HPwSmFP88niiJg.mp4" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Interactive progress across parallel engines. Those progress bars are interactive widgets running locally, &lt;em&gt;and on each remote engine!&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;For a lot of today’s workloads, my default recommendation is: use &lt;a href="https://dask.org/"&gt;dask&lt;/a&gt; or bodo or another modern tool. IPython even makes this easier if you already happen to have an IPython Parallel cluster, you can tell it to “&lt;a href="https://ipyparallel.readthedocs.io/en/8.0.0/examples/dask.html"&gt;become dask&lt;/a&gt;,” and off you go:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;dask_client = rc.become_dask(ncores=1)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;However, where IPython Parallel can shine, is making traditional &lt;a href="https://en.wikipedia.org/wiki/SPMD"&gt;SPMD&lt;/a&gt; (e.g., MPI) workloads interactive, especially for prototyping and debugging.&lt;/p&gt;
&lt;p&gt;If you have an MPI simulation and you wish you could pause it in the middle, poke around and make plots and interact with just one node or all of them with all the interactive tools available to you in Jupyter, IPython Parallel may be the tool for you.&lt;/p&gt;
&lt;h2 id="recent-developments"&gt;Recent developments&lt;/h2&gt;
&lt;p&gt;Because I see the main problem IPython Parallel solves well is adding interactivity to direct parallel execution, the focus of recent developments has been on improving that story, based on feedback from users who are often in traditional HPC environments like SLURM or PBS, or using MPI in the cloud.&lt;/p&gt;
&lt;p&gt;Much of the feedback on challenges over the years have been around:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Scaling to larger numbers of engines (1,–10,000)&lt;/li&gt;
&lt;li&gt;Better feedback and recovery when things go wrong&lt;/li&gt;
&lt;li&gt;Security requirements&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Thanks to &lt;a href="https://bodo.ai/"&gt;Bodo&lt;/a&gt;, I’ve been able to spend a lot of time this year addressing several of those issues, now available as IPython Parallel 8.0.&lt;/p&gt;
&lt;p&gt;Last year, Tom-Olav Bøyum developed a broadcast scheduler as part of his &lt;a href="https://www.duo.uio.no/handle/10852/81515"&gt;Master’s thesis&lt;/a&gt;, which vastly improves the efficiency of sending the same task to all engines (the main thing we do in IPP+MPI). This work is now mostly done and landed in IPython Parallel 7. We’ve also addressed longstanding issues with registering large numbers of engines (5–10,000) at once, a common source of failure on HPC clusters.&lt;/p&gt;
&lt;p&gt;One of the most frustrating / challenging things when working with MPI or other parallel computations is dealing with hangs and restarting tasks. Through the new &lt;a href="https://ipyparallel.readthedocs.io/en/8.0.0/examples/Cluster%20API.html"&gt;Cluster API&lt;/a&gt;, IPython Parallel now has one of its most requested features: the ability to send signals to one or all engines, and forcefully restart the cluster if it’s stuck.&lt;/p&gt;
&lt;p&gt;We’ve also improved interactive feedback with progress bars, live streaming output, a new JupyterLab plugin based on dask-labextension, and support for the Jupyter widget protocol, so you can instantiate widgets on engines and interact with them directly from a notebook.&lt;/p&gt;
&lt;p&gt;Security questions have also been raised, because IPython (and Jupyter’s) use of ZeroMQ assumes a ‘trusted network’, typically localhost or BSD sockets for notebooks. Since IPython Parallel usually runs across a network, security can be more of a concern (whereas notebooks typically only use standard HTTPS connections over a network), and the level of security in the Jupyter protocol alone may not be sufficient. In the past, we’ve resorted to tunneling TCP over SSH for more secure network traffic, but still always implicitly trusting localhost.&lt;/p&gt;
&lt;p&gt;IPython Parallel 7.1 enables &lt;a href="https://rfc.zeromq.org/spec/26/"&gt;CurveZMQ&lt;/a&gt; for full authentication, encryption, and &lt;a href="https://en.wikipedia.org/wiki/Forward_secrecy"&gt;forward-secrecy&lt;/a&gt; at the transport level, solving a longstanding security shortcoming of IPython Parallel.&lt;/p&gt;
&lt;h2 id="ipython-parallel-in-the-future"&gt;IPython Parallel in the future&lt;/h2&gt;
&lt;p&gt;Scaling is the biggest challenge for IPython Parallel due to its communication model. However, that scaling is on the order of IPython &lt;em&gt;engines&lt;/em&gt;, which is not the same as the number of cores. You can benefit from this today if your tasks are already multi-threaded, e.g., through OpenMP threads in numpy, or your own Python threads or multiprocess breakdown of tasks.&lt;/p&gt;
&lt;p&gt;Better support of multi-level parallelism, where each “engine” may represent a multicore node, as individual nodes get bigger and bigger would mean moving the bar where IPython Parallel communication is the bottleneck back two orders of magnitude on large machines, because one ‘engine’ could represent 128 cores or more. Good support for 1000 128-core nodes would give IPython Parallel some pretty comfortable headroom for our target use cases, scale-wise.&lt;/p&gt;
&lt;p&gt;Also related to scaling, bringing the BroadcastView to maturity should allow us to completely replace the DirectView scheduler, as it should be able to match or beat it in almost every scenario with some further development. There are some features lacking, especially when it comes to error handling, but those can certainly be addressed in time.&lt;/p&gt;
&lt;p&gt;Finally, I’ll invite you to get involved. Especially if you are interested in prototyping parallel code, or making traditional MPI-style code interactive, check out IPython Parallel 8 and &lt;a href="https://github.com/ipython/ipyparallel/issues"&gt;let us know&lt;/a&gt; how it goes, or contribute your use case as an example.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;pip install --upgrade ipyparallel
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
</content><category term="HPC"/><category term="IPython"/></entry><entry><title>National Scale Interactive Computing</title><link href="https://blog.jupyter.org/posts/2019/national-scale-interactive-computing/" rel="alternate"/><published>2019-08-22T19:06:00+00:00</published><updated>2019-08-22T20:06:00+00:00</updated><author><name>James Colliander</name></author><id>tag:blog.jupyter.org,2019-08-22:/posts/2019/national-scale-interactive-computing/</id><summary type="html">&lt;p&gt;Delivering interactive computing to universities at a national scale with a Jupyter stack.&lt;/p&gt;
</summary><content type="html">&lt;p&gt;&lt;em&gt;This is an invited post from Jim Colliander, Professor of Mathematics at UBC and Director of the &lt;a href="http://www.pims.math.ca/"&gt;Pacific Institute for the Mathematical Sciences&lt;/a&gt;.&lt;/em&gt;&lt;sup class="footnote-ref"&gt;&lt;a href="#fn1" id="fnref1"&gt;[1]&lt;/a&gt;&lt;/sup&gt;&lt;/p&gt;
&lt;p&gt;In 2017, the &lt;a href="http://www.pims.math.ca/"&gt;Pacific Institute for the Mathematical Sciences (PIMS)&lt;/a&gt;, in partnership with &lt;a href="https://www.computecanada.ca/featured/compute-canada-and-pims-launch-jupyter-service-for-researchers/"&gt;Compute Canada&lt;/a&gt; and &lt;a href="https://www.cybera.ca/services/jupyter-all-in-one-science-platform/"&gt;Cybera&lt;/a&gt;, launched &lt;a href="https://syzygy.ca"&gt;Syzygy&lt;/a&gt;, a cloud-hosted interactive computing platform that delivers &lt;a href="https://jupyter.org/"&gt;JupyterHub deployments&lt;/a&gt; for &lt;a href="https://www.google.com/maps/d/embed?mid=1nzSAGLSn8eWdfQ6K7zTw-31h82I&amp;amp;hl=en"&gt;universities across Canada&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Syzygy has been used by over 16,000 students at 20 universities. The main results of the Syzygy experiment so far are:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Demand for interactive computing is ubiquitous&lt;sup class="footnote-ref"&gt;&lt;a href="#fn2" id="fnref2"&gt;[2]&lt;/a&gt;&lt;/sup&gt; and growing strongly at universities.&lt;/li&gt;
&lt;li&gt;The Jupyter ecosystem is an effective way to deliver interactive computing.&lt;/li&gt;
&lt;li&gt;A scalable, sustainable, and cost-effective interactive computing service for universities is needed as soon as possible.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="demand-for-interactive-computing"&gt;Demand for interactive computing&lt;/h2&gt;
&lt;p&gt;Both research and teaching at universities are adapting to major societal changes driven by explosions in data and computational tools. New educational programs that prepare students to think computationally are emerging, while research strategies are changing in ways that are more open, reproducible, collaborative, and interdisciplinary. These transformations are inextricably linked and are accelerating demand for interactive computing. The Syzygy experiment has shown that using Jupyter in educational programs drives interest in using Jupyter for research (and vice versa). For example, students in mathematics, statistics, and computer science &lt;a href="https://medium.com/pims-math/saving-lives-with-data-and-math-b697667d1cd7"&gt;collaborated with a researcher from St. Paul’s Hospital in Vancouver using Syzygy&lt;/a&gt; to identify new pathways to prevent death from sepsis. Research communities typically need access to deeper computational resources and often span multiple universities, but the common thread is the need to expand access to interactive computing.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="A map of JupyterHub deployments deployed by Syzygy." src="https://blog.jupyter.org/posts/2019/national-scale-interactive-computing/images/001-1_L8MzmheO2NZQBH0t-BGpFg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;A map of JupyterHub deployments deployed by Syzygy.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="technical-milestone-achieved"&gt;Technical milestone achieved&lt;/h2&gt;
&lt;p&gt;The Syzygy project has demonstrated that it’s possible to deploy tools for interactive computation at a national scale rapidly and efficiently using an entirely open source technology stack. Students, faculty and staff across Canada use Syzygy to access Jupyter through their browsers with their university single-sign-on credentials. The JupyterHubs range from a “standard” configuration to bespoke environments with specially curated tools and data integrations. This richness is possible because of the architecture of the Syzygy and Jupyter projects and the flexibility of the underlying cloud resources. As a case-study, Syzygy demonstrates that the Jupyter community has achieved a significant technical milestone: interactive computing &lt;em&gt;can be delivered&lt;/em&gt; at national scale using cloud technologies.&lt;/p&gt;
&lt;h2 id="service-level-requirements"&gt;Service level requirements&lt;/h2&gt;
&lt;p&gt;The validation that interactive computing can be technically delivered at national scale prompts universities to ask a variety of questions. Can interactive computing service be delivered robustly? How will users be supported? What are the uptime expectations? What is the data security policy? How is privacy protected? Can the robustness of the service be clarified in a service level agreement? Syzygy, as an experimental service offered to universities at no charge and without a service level agreement, does not properly address these questions. To advance on their education-research-service mission and address growing demand, universities need a reliable interactive computing service with a service level agreement.&lt;/p&gt;
&lt;h2 id="whats-next"&gt;What’s next?&lt;/h2&gt;
&lt;p&gt;How should universities address their needs for interactive computing over the next five years? Right now, universities are following two primary approaches:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;🙏 &lt;em&gt;&lt;strong&gt;Ad hoc&lt;/strong&gt;&lt;/em&gt;: faculty figure out how to meet their own needs for interactive computing; IT staff deploys JupyterHub on local or commercial cloud servers; this approach gives universities control over their deployments and hardware, though requires time and expertise that many may not have.&lt;/li&gt;
&lt;li&gt;🎩 &lt;em&gt;&lt;strong&gt;Use a cloud provider’s service&lt;/strong&gt;&lt;/em&gt;: Google Colab, Amazon Sagemaker, Microsoft Azure Notebooks, IBM Watson Studio; this approach allows universities to quickly launch interactive computing services, with a loss of flexibility and some risks by becoming reliant upon a particular vendor’s closed-source and proprietary software.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These approaches are not sustainable over the long term. If universities all deploy their own JupyterHub services, many will need technical expertise they do not currently have and will involve a significant duplication of effort. If universities rely on hosted cloud notebook services, the reliance on proprietary technology will impair their ability to switch between different cloud vendors, change hardware, customize software, etc. Vendor lock-in will limit the ability of universities to respond to changes in price for the service. Universities will lose agility in responding to changes in faculty, staff, and student computing needs.&lt;/p&gt;
&lt;p&gt;There is a third option that addresses the issues with these two approaches and generates other benefits for universities:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;🤔 &lt;em&gt;&lt;strong&gt;Form an interactive computing consortium&lt;/strong&gt;&lt;/em&gt;: universities collaborate to &lt;em&gt;build&lt;/em&gt; an interactive computing service provider aligned with their missions to better serve their students, facilitate research, and avoid risks associated with vendor lock-in.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;To retain control over their interactive computing stacks, avoid dependence&lt;sup class="footnote-ref"&gt;&lt;a href="#fn3" id="fnref3"&gt;[3]&lt;/a&gt;&lt;/sup&gt; on cloud providers, and accelerate the emergence of new programs, universities should work together to deploy interactive computing environments in a vendor-agnostic manner. This might take the form of a consortium — an organization dedicated to serving the needs of universities through customized shared infrastructure for interactive computing. The consortium would also ensure that universities will continue to play a leadership role in the development of the interactive computing tools used for education and research.&lt;/p&gt;
&lt;p&gt;The Syzygy experiment confirmed that growing demand for interactive computation within universities can be supplied with the available technologies advanced by the Jupyter open source community. In the coming year, we aim to build upon the success of the Syzygy experiment and seed an initial node of a consortium in Canada with the intention of fostering a global network of people invested in advanced interactive computing. If you are interesting in partnering, &lt;a href="https://ten.blue/2i2c/#/3/4"&gt;please get in touch!&lt;/a&gt; See &lt;a href="https://discourse.jupyter.org/t/creating-national-infrastructure-for-jupyter-environments/1966"&gt;this Jupyter Community Forum post&lt;/a&gt; to continue the discussion.&lt;/p&gt;
&lt;hr class="footnotes-sep"&gt;
&lt;section class="footnotes"&gt;
&lt;ol class="footnotes-list"&gt;
&lt;li id="fn1" class="footnote-item"&gt;&lt;p&gt;The author gratefully acknowledges feedback on this piece from Ian Allison, Lindsey Heagy, Chris Holdgraf, Fernando Perez, and Lindsay Sill. &lt;a href="#fnref1" class="footnote-backref"&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn2" class="footnote-item"&gt;&lt;p&gt;Interactive computing needs have been identified in agriculture, applied mathematics, astronomy, chemistry, climate science, computer science, data science, digital humanities, ecology, economics, engineering, genomics, geoscience, health sciences, K-12 education, neuroscience, political science, physics, pure mathematics, statistics, and sociology. &lt;a href="#fnref2" class="footnote-backref"&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn3" class="footnote-item"&gt;&lt;p&gt;Relying on commercial cloud vendors to provide the interactive computing service for universities risks recreating the problems associated with scientific publishing that emerged with the internet. &lt;a href="#fnref3" class="footnote-backref"&gt;↩︎&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/section&gt;
</content><category term="cloud computing"/><category term="education"/><category term="HPC"/><category term="JupyterHub"/><category term="science"/></entry><entry><title>Jupyter for Science User Facilities and High Performance Computing</title><link href="https://blog.jupyter.org/posts/2019/jupyter-for-science-user-facilities-and-high/" rel="alternate"/><published>2019-07-09T23:53:00+00:00</published><updated>2019-07-09T23:53:00+00:00</updated><author><name>Rollin Thomas</name></author><id>tag:blog.jupyter.org,2019-07-09:/posts/2019/jupyter-for-science-user-facilities-and-high/</id><summary type="html">&lt;p&gt;Jupyter is the “Google Docs” of data science. It provides that same kind of easy-to-use ecosystem, but for interactive data exploration, modeling, and analysis.&lt;/p&gt;</summary><content type="html">&lt;p&gt;Jupyter is the &lt;a href="https://www.nature.com/articles/d41586-018-07196-1"&gt;“Google Docs” of data science&lt;/a&gt;. It provides that same kind of easy-to-use ecosystem, but for interactive data exploration, modeling, and analysis. Just as people have come to expect to be able to use Google Docs everywhere, scientists assume that Jupyter is there for them whenever and wherever they open their laptops.&lt;/p&gt;
&lt;p&gt;But what if the data you want to interact with through Jupyter doesn’t fit on your laptop or is excruciating to move? What if the model you want to build and test requires more computing power and storage than you have right in front of you? As a scientist, you want the same interactive experience and all the benefits of Jupyter, but you also need to “reach out” to put something big into your science process: A supercomputer, a telescope data archive, a beam-line at a synchrotron. Can Jupyter help you do that big science? What efforts are in motion already to make this a reality, what work still needs to be done, and who needs to do it?&lt;/p&gt;
&lt;p&gt;Doing this right will take a community: New collaborations between core Jupyter developers, engineers from high-performance computing (HPC) centers, staff from large-scale experimental and observational data (EOD) facilities, users and other stakeholders. Many facilities have figured out how to deploy, manage, and customize Jupyter, but have done it while focused on their unique requirements and capabilities. Still others are just taking their first steps and want to avoid reinventing the wheel. With some initial critical mass, we can start contributing what we’ve learned separately into a shared body of knowledge, patterns, tools, and best practices.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="40+ participants from universities, national labs, industry, and science user facilities. Credit: Fernando Perez." src="https://blog.jupyter.org/posts/2019/jupyter-for-science-user-facilities-and-high/images/001-1_VqdM1ZzoT6oepd6UZCRIyA.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;40+ participants from universities, national labs, industry, and science user facilities. Credit: Fernando Perez.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;In June, a &lt;a href="https://blog.jupyter.org/posts/2019/jupyter-community-workshop-jupyter-for-scientific-user/"&gt;Jupyter Community Workshop&lt;/a&gt; held at the National Energy Research Scientific Computing Center (NERSC) and the Berkeley Institute for Data Science (BIDS) brought about 40 members of this community together to start distilling. Over &lt;a href="https://jupyter-workshop-2019.lbl.gov/agenda"&gt;three days&lt;/a&gt; in talks and breakout sessions, we addressed pain points and best practices in Jupyter deployment, infrastructure, and user support; securing Jupyter in multi-tenant environments; sharing notebooks; HPC/EOD-focused Jupyter extensions; and strategies for communication with stakeholders.&lt;/p&gt;
&lt;p&gt;Here are just a few highlights from the meeting:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Michael Milligan from the Minnesota Supercomputing Center perfectly set the tone for the workshop with his keynote, &lt;a href="https://drive.google.com/a/lbl.gov/file/d/1YXuwwHSM1NqUKBkutv1YrJ3Fzsj2UnFN/view?usp=sharing"&gt;“Jupyter is a One-Stop Shop for Interactive HPC Services.”&lt;/a&gt; Michael is the creator of &lt;a href="https://github.com/jupyterhub/batchspawner"&gt;BatchSpawner&lt;/a&gt; and &lt;a href="https://github.com/jupyterhub/wrapspawner"&gt;WrapSpawner&lt;/a&gt;, JupyterHub Spawners that let HPC users run notebooks on compute nodes supporting a variety of batch queue systems. Contributors to both packages met in an afternoon-long breakout to build consensus around some technical issues, start managing development and support in a collaborative way, and gel as a team.&lt;/li&gt;
&lt;li&gt;Securing Jupyter is a huge topic. Thomas Mendoza from Lawrence Livermore National Laboratory talked about &lt;a href="https://github.com/jupyterhub/jupyterhub/pull/2055"&gt;his work&lt;/a&gt; to enable &lt;a href="https://drive.google.com/file/d/16N44SPtKZyPKlcDWp8G_mJcQq-g_G0e2/view"&gt;end-to-end SSL in JupyterHub and best practices for securing Jupyter&lt;/a&gt;. Outcomes from two breakouts on security include a plan to more prominently document security best practices, and a future meeting (perhaps another Jupyter Community Workshop?) focused specifically on security in Jupyter.&lt;/li&gt;
&lt;li&gt;Speakers from Lawrence Livermore and Oak Ridge National Laboratories, the European Space Agency showed off a variety of beautiful JupyterLab extensions, integrations, and plug-ins for climate science, complex physical simulations, astronomical images and catalogs, and atmospheric monitoring. People at a variety of facilities are finding ways to adapt Jupyter to meet the specific needs of their scientists.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Really, there’s just too much to pack into a blog post so we encourage you to look at the &lt;a href="https://jupyter-workshop-2019.lbl.gov/agenda"&gt;talk slides&lt;/a&gt; and &lt;a href="https://discourse.jupyter.org/t/notes-from-breakout-sessions/1338"&gt;notes on Discourse&lt;/a&gt; — all the breakout notes have been posted there to &lt;a href="https://discourse.jupyter.org/c/jupyterhub/hpc-meeting-2019"&gt;this topic&lt;/a&gt;. We’re working on getting videos of the slide presentations up on the workshop website as well. Watch for announcements of future meeting opportunities and documentation on Discourse as well.&lt;/p&gt;
&lt;p&gt;Finally we want to thank Project Jupyter, NumFOCUS, and Bloomberg for their help making this meeting happen. We all came away with a better sense of who is doing what in our community, and how we can work together on this new area of growth for the Jupyter community. The organizers also want to thank their respective institutions’ administrative staff (Seleste Rodriguez at NERSC, and Stacy Dorton at BIDS) for helping with workshop logistics.&lt;/p&gt;
</content><category term="events"/><category term="HPC"/><category term="JupyterHub"/><category term="science"/><category term="workshops"/></entry><entry><title>Jupyter Community Workshop: Jupyter for Scientific User Facilities and High-Performance Computing</title><link href="https://blog.jupyter.org/posts/2019/jupyter-community-workshop-jupyter-for-scientific-user/" rel="alternate"/><published>2019-01-29T19:35:00+00:00</published><updated>2019-01-29T19:35:00+00:00</updated><author><name>Rollin Thomas</name></author><id>tag:blog.jupyter.org,2019-01-29:/posts/2019/jupyter-community-workshop-jupyter-for-scientific-user/</id><summary type="html">&lt;p&gt;We are excited to share more news about the Jupyter Community Workshop for Scientific User Facilities and High-Performance Computing!&lt;/p&gt;</summary><content type="html">&lt;p&gt;We are excited to share more news about the Jupyter Community Workshop for Scientific User Facilities and High-Performance Computing! This is part of a series of &lt;a href="https://blog.jupyter.org/posts/2019/jupyter-community-workshops/"&gt;Jupyter Community Workshops&lt;/a&gt; funded by &lt;a href="https://www.techatbloomberg.com/"&gt;Bloomberg&lt;/a&gt; to “bring together small groups of Jupyter community members and core contributors for high-impact strategic work and community engagement on focused topics.”&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;This workshop will be held in Berkeley, California, from Tuesday June 11 to Thursday June 13, 2019&lt;/strong&gt;. The workshop is being jointly hosted at the National Energy Research Scientific Computing Center (&lt;a href="http://www.nersc.gov/"&gt;NERSC&lt;/a&gt;, part of &lt;a href="https://www.lbl.gov/"&gt;Lawrence Berkeley National Laboratory&lt;/a&gt;) and the Berkeley Institute for Data Science (&lt;a href="https://bids.berkeley.edu/"&gt;BIDS&lt;/a&gt;, at the &lt;a href="https://www.berkeley.edu/"&gt;University of California&lt;/a&gt;). The workshop executive committee consists of Rollin Thomas (NERSC), Dan Allan (Brookhaven National Laboratory), and Chris Holdgraf (BIDS — UC Berkeley).&lt;/p&gt;
&lt;p&gt;We, the organizers, invite your expression of interest in the effort through &lt;a href="https://docs.google.com/forms/d/e/1FAIpQLSdAoyJ6Hub36XgLXnJDbgGwVbVCSgzdn5X-NPsOVph7nzJP9Q/viewform?usp=sf_link"&gt;&lt;em&gt;&lt;strong&gt;this Google form&lt;/strong&gt;&lt;/em&gt;&lt;/a&gt;. Let us know whether you’d be interested potentially in attending, just want to be kept in the loop, or just want to express your support. Finding out who is doing what with Jupyter in this space is the first step in building our community.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Why?&lt;/strong&gt; Advances in technology at experimental and observational science facilities (EOS facilities: telescopes, particle accelerators, light sources, genome sequencers and so on), in robust high-bandwidth global networks, and in high-performance computing (HPC) have resulted in an exponential growth of data for scientists to collect, manage, and understand. Interpreting these data streams requires computational and storage resources greatly exceeding those available on laptops, workstations, or university department clusters. Funding agencies increasingly look to HPC centers to address the growing and changing data needs of their scientists. These institutions are uniquely equipped to provide the resources needed for extreme scale science. At the same time, scientists seek new ways to seamlessly and transparently integrate HPC into their EOS workflows. That’s where we see Jupyter fitting in.&lt;/p&gt;
&lt;p&gt;We know that scientists love Jupyter because it combines visualization, data analytics, text, and code into a document they can share, modify, and even publish. What about using Jupyter to control experiments in real-time, or steer complex simulations on a supercomputer, or even connect experiments to HPC for real-time feedback and decision making? How can users reach outside the notebook to corral external data and computational resources in a seamless, Jupyter-friendly manner?&lt;/p&gt;
&lt;p&gt;These were the questions on our minds when we proposed a three-day workshop for Jupyter developers, HPC engineers, and staff from EOS facilities. We are looking to foster a new collaborative community that can make Jupyter the pre-eminent interface for managing EOS workflows and data analytics at HPC centers. EOS scientists need Jupyter to work well at their facilities and HPC centers, and this workshop will help us address the technical, sociological, and policy challenges involved.&lt;/p&gt;
&lt;p&gt;The workshop itself will include presentations, posters, and a couple half-day hack-a-thon/breakout sessions for collaboration. We will identify best practices, share lessons learned, clarify gaps and challenges in supporting deployments, and work on new tools to make Jupyter easier to use for big science.&lt;/p&gt;
&lt;p&gt;During the workshop, participants will be invited to collaborate on a survey white paper that documents the current state of the art in Jupyter deployments at various facilities and HPC centers. The document will include deployment descriptions, maintenance and user support strategies, security discussions, use cases, and lessons learned. A forward-looking summary provided at the end of the white paper will tie together common threads across various facilities and highlight areas for future research, development, and implementation. We will aim to have the paper completed and published to arXiv within three months of the end of the workshop. These ideas in a single document should help developers, maintainers, and researchers make the case they need to management and policymakers to drive the effort forward.&lt;/p&gt;
&lt;p&gt;So let us know if you’re interested in the effort by filling out &lt;a href="https://docs.google.com/forms/d/e/1FAIpQLSdAoyJ6Hub36XgLXnJDbgGwVbVCSgzdn5X-NPsOVph7nzJP9Q/viewform"&gt;the Google form&lt;/a&gt;, even if you think you can’t make it. Part of what we’re doing is finding out who is doing what with Jupyter where in HPC and EOS. That’s the real first step in building our community!&lt;/p&gt;
</content><category term="events"/><category term="HPC"/><category term="science"/><category term="workshops"/></entry></feed>