Geophysics in the Cloud Competition

Join the 2021 GSH Geophysics in the cloud competition. Build a novel seismic inversion app and access all the data on demand with serverless cloud storage. Example notebooks show how to access this data and use AWS SageMaker to build your ML models. With prizes.

Author: Ben Lasscock, Ph.D.

Geophysics in the Cloud Competition

The 2021 Houston GSH Geophysics on Cloud Competition is sponsored by AWS Energy and Enthought. This competition will allow teams and individuals to develop new and innovative solutions for seismic inversion. We’re going all in on cloud. You will be provided the latest technologies for serverless access to big data, examples of AWS SageMaker to learn how to build ML models on the cloud and a gather.town, where we can work and collaborate in 8-bit.

Access All the Data

A common theme when discussing AI/ML in exploration geophysics has been that only a very small percentage of available data is used in analysis and decision making. One of the goals of this competition is to make ALL data available to the participants, on demand.

This competition presents both a logistical and technical challenge for both organizers and participants. The seismic datasets are large. Downloading this data would typically take hours, a cost multiplied across each and every participant. For the organizers, we don’t want to see the work of loading and manipulating large SEGY format data replicated across the teams. More overhead loading data means less time (and less fun) developing ML for seismic inversion.

While we want participants to have access to ALL the data, we expect they will only use what they find to be relevant in solving the competition problem. This detail is important when using specialist GPUs and tools like AWS SageMaker to build models. We don’t want to be wasting valuable compute time doing I/O.

Going Serverless

Competition datasets will be made available to the participants through a convenient api, the data reformatted for efficient serverless access. Serverless means that the data can be accessed directly from blob storage (S3). For the organizers, we don’t have to manage an extra server to provide access to data. For the participants, it means efficient access to the parts of the dataset they want, on demand.

One such efficient format of seismic data is OpenVDS. OpenVDS provides fast access to slices (inline, crossline, and time) and 3D chunks. The upcoming release of OpenVDS+, by Bluware, provides an easy (pip installable) library that participants can use in their notebooks. OpenVDS is also part of the OSDU Data Platform, so we should be seeing a lot more of it in the future.

Get Started

The problem of assembling an AI/ML ready data set has been solved by using a serverless model, making the most of the scarce resources available for the competition.

This story really isn’t too different from what we see in the industry at large: how to get the most innovation with the least expenditure while making highly efficient use of expert time.

Let the competition begin. Entries close 26 March, and the competition begins 1 April. No foolin’.

Visit the website to learn more and enter the competition.

About the Author

Ben Lasscock, holds a Ph.D. and a B.Sc. in theoretical physics as well as a B.Sc. in physics and theoretical physics from the University of Adelaide. Before coming to geoscience, Ben worked as a portfolio manager at a large hedge fund in Australia. He has publications in the areas of high energy physics, Bayesian time series analysis and geophysics.

Share this article:

Related Content

Prospecting for Data on the Web

Introduction At Enthought we teach a lot of scientists and engineers about using Python and the ecosystem of scientific Python packages for processing, analyzing, and…

Read More

Configuring a Neural Network Output Layer

Introduction If you have used TensorFlow before, you know how easy it is to create a simple neural network model using the Keras API. Just…

Read More

No Zero Padding with strftime()

One of the best features of Python is that it is platform independent. You can write code on Linux, Windows, and MacOS and it works…

Read More

Got Data?

Introduction So, you have data and want to get started with machine learning. You’ve heard that machine learning will help you make sense of that…

Read More

Sorting Out .sort() and sorted()

Sorting Out .sort() and sorted() Sometimes sorting a Python list can make it mysteriously disappear.  This happens even to experienced Python programmers who use .sort()…

Read More

A Beginner’s Guide to Deep Learning

Deep learning. By this point, we’ve all heard of it. It’s the magic silver bullet that can fix any complex problem. It’s the special ingredient…

Read More

Giving Visibility to Renewable Energy

The ultimate project goal of EnergizAIR Infrastructure was to raise individual awareness of the contribution of renewable energy sources, and ultimately change behaviors. Now ten…

Read More

Introducing Enthought Edge: Unlocking the Value of R&D Data

While the value of R&D data is clear, finding a way to sort through it can be daunting given the special handling required to extract…

Read More

Machine Learning in Materials Science

The process of materials discovery is complex and iterative, requiring a level of expertise to be done effectively. Materials workflows that require human judgement present…

Read More

AI Needs the ‘Applied Sciences’ Treatment

As industries rapidly advance in AI/machine learning, a key to unlocking the power of these approaches for companies is an enabling environment. Domain experts need…

Read More

Join Our Mailing List!

Sign up below to receive email updates including the latest news, insights, and case studies from our team.