menu
arrow_back

Distributed Image Processing in Cloud Dataproc

Distributed Image Processing in Cloud Dataproc

1 jam 7 Kredit

GSP010

Google Cloud Self-Paced Labs

Overview

In this hands-on lab, you will learn how to use Apache Spark on Cloud Dataproc to distribute a computationally intensive image processing task onto a cluster of machines. This lab is part of a series of labs on processing scientific data.

What you'll learn

  • How to create a managed Cloud Dataproc cluster with Apache Spark pre-installed.

  • How to build and run jobs that use external packages that aren't already installed on your cluster.

  • How to shut down your cluster.

Prerequisites

This is an advanced level lab. Familiarity with Cloud Dataproc and Apache Spark is recommended, but not required. If you're looking to get up to speed in these services, be sure to check out the following labs:

Once you're ready, scroll down to learn more about the services that you'll be using in this lab.

Bergabunglah dengan Qwiklabs untuk membaca tentang lab ini selengkapnya... beserta informasi lainnya!

  • Dapatkan akses sementara ke Google Cloud Console.
  • Lebih dari 200 lab mulai dari tingkat pemula hingga lanjutan.
  • Berdurasi singkat, jadi Anda dapat belajar dengan santai.
Bergabung untuk Memulai Lab Ini
Skor

—/30

Create a development machine in Compute Engine

Jalankan Langkah

/ 5

Install Software in the development machine

Jalankan Langkah

/ 5

Create a GCS bucket

Jalankan Langkah

/ 5

Download some sample images into your bucket

Jalankan Langkah

/ 5

Create a Cloud Dataproc cluster

Jalankan Langkah

/ 5

Submit your job to Cloud Dataproc

Jalankan Langkah

/ 5