Category: Big Data
New – Upload AWS Cost & Usage Reports to Redshift and QuickSight
Many AWS customers have been asking us for a way to programmatically analyze their Cost and Usage Reports (read New – AWS Cost and Usage Reports for Comprehensive and Customizable Reporting for more info). These customers are often using AWS to run multiple lines of business, making use of a wide variety of services, often spread out across multiple regions. Because we provide very detailed billing and cost information, this is a Big Data problem and one that can be easily addressed using AWS services!
While I was on vacation earlier this month, we launched a new feature that allows you to upload your Cost and Usage reports to Amazon Redshift and Amazon QuickSight. Now that I am caught up, I’d like to tell you about this feature.
Upload to Redshift
I started by creating a new Redshift cluster (if you already have a running cluster, you need not create another one). Here’s my cluster:

Next, I verified that I had enabled the Billing Reports feature:

Then I hopped over to the Cost and Billing Reports and clicked on Create report:

Next, I named my report (MyReportRedshift), made it Hourly, and enabled support for both Redshift and QuickSight:

I wrapped things up by selecting my delivery options:

I confirmed my desire to create a report on the next page, and then clicked on Review and Complete. The report was created and I was informed that the first report would arrive in the bucket within 24 hours:

While I was waiting I installed PostgreSQL on my EC2 instance (sudo yum install postgresql94) and verified that I was signed up for the Amazon QuickSight preview. Also, following the directions in Create an IAM Role, I made a read-only IAM role and captured its ARN:

Back in the Redshift console, I clicked on Manage IAM Roles and associated the ARN with my Redshift cluster:

The next day, I verified that the files were arriving in my bucket as expected, and then returned to the console in order to retrieve a helper file so that I could access Redshift:

I clicked on Redshift file and then copied the SQL command:

I inserted the ARN and the S3 region name into the SQL (I had to add quotes around the region name in order to make the query work as expected):

And then I connected to Redshift using psql (I can use any visual or CLI-based SQL client):
$ psql -h jbcluster.XYZ.us-east-1.redshift.amazonaws.com \
-U root -p 5439 -d dev
Then I ran the SQL command. It created a pair of tables and imported the billing data from S3.
Querying Data in Redshift
Using some queries supplied by my colleagues as a starting point, I summed up my S3 usage for the month:

And then I looked at my costs on a per-AZ basis:

And on a per-AZ, per-service basis:
Just for fun, I spent some time examining the Redshift Console. I was able to see all of my queries:

Analyzing Data with QuickSight
I also spent some time analyzing the cost and billing data using Amazon QuickSight. I signed in and clicked on Connect to another data source or upload a file:

Then I dug in to my S3 bucket (jbarr-bcm) and captured the URL of the manifest file (MyReportRedshift-RedshiftManifest.json):

I selected S3 as my data source and entered the URL:

QuickSight imported the data in a few seconds and the new data source was available. I loaded it into SPICE (QuickSight’s in-memory calculation engine). With three or four more clicks I focused on the per-AZ data, and excluded the data that was not specific to an AZ:

Another click and I switched to a pie chart view:

I also examined the costs on a per-service basis:

As you can see, the new data and the analytical capabilities of QuickSight allow me (and you) to dive deep into your AWS costs in minutes.
Available Now
This new feature is available now and you can start using it today!
— Jeff;
Learn about Amazon Redshift in our new Data Warehousing on AWS Class
As our customers continue to look to use their data to help drive their missions forward, finding a way to simply and cost-effectively make use of analytics is becoming increasingly important. That is why I am happy to announce the upcoming availability of Data Warehousing on AWS, a new course that helps customers leverage the AWS Cloud as a platform for data warehousing solutions.
New Course
Data Warehousing on AWS is a new three-day course that is designed for database architects, database administrators, database developers, and data analysts/scientists. It introduces you to concepts, strategies, and best practices for designing a cloud-based data warehousing solution using Amazon Redshift. This course demonstrates how to collect, store, and prepare data for the data warehouse by using other AWS services such as Amazon DynamoDB, Amazon EMR, Amazon Kinesis, and Amazon S3. Additionally, this course demonstrates how you can use business intelligence tools to perform analysis on your data. Organizations who are looking to get more out of their data by implementing a Data Warehousing solution or expanding their current Data Warehousing practice are encouraged to sign up.
These classes (and many more) are available through AWS and our Trainintg Partners. Find upcoming classes in our global training schedule or learn more at AWS Training.
— Jeff;
New AWS Digital Library for Big Data Solutions
My colleague Luis Daniel Soto has been working with AWS Community Hero Lynn Langit to create a comprehensive collection of resources for customers who are ready to run Big Data applications on AWS!
Here’s what they have to say….
— Jeff;
Today the AWS Marketplace is launching a new on-line video library designed to help our customers find AWS Marketplace vendor solutions, as well as accelerate and manage short and long-term data integration, business intelligence and advanced analytics projects for their AWS cloud and on-premises data.
The AWS Marketplace Digital Library for Big Data provides business and technical content from AWS Marketplace technology vendors and case studies from customers who have built end-to-end Big Data solutions. The segments are hosted by cloud and Big Data architect Lynn Langit and organized around a common set of functionality to help organizations and individuals find the AWS Marketplace vendor solutions to address their particular needs.
The library is hosted on a video webcasting platform which allows our customers to interact with AWS Marketplace partners, by asking questions as they watch the demos and interviews in split-screen mode. Here’s a sample:

If you are an APN Partner and want to learn more or want to be part of the AWS Digital Library, visit the new Big Data Partner Solutions page.
— Luis and Lynn
New – AstroCompute in the Cloud Grants Program
The skills and techniques needed to create, store, process, and manage data sets that start in the hundreds of gigabytes and grow to multiple terabyte size are all too rare. It is time to change that!
We have teamed up with the Square Kilometre Array (SKA) to create the new AstroCompute in the Cloud grant program in order to address these “to infinity and beyond” sorts of problems in order to ensure that mature, high-quality data management and processing solutions are in place by the time the SKA starts to pump out data in 2020 or so.
What’s the SKA?
I should start with a quick introduction to the SKA. It is funded by 11 world governments—with more planning to join—and will have a physical footprint in South Africa and Australia. The site in South Africa will be home to a set of high and mid frequency dish antennas (200 to start, with plans to grow to 2000 over time, spread across eight other African countries). The site in Australia will host 125,000 low frequency dipole antennas at the start, growing to a total of one million by the late 2020’s. These antennas will allow astronomers to monitor the sky in unprecedented detail and to run whole-sky surveys far faster than any system currently in existence. The goal is to address—and hopefully tackle—some of the most fundamental questions about our universe.
All of this raw data (exabytes per day at full scale) will be distilled down to a far more manageable level (exabytes per year) for storage and analysis. The distillation (filtering, calibration, geometric transformation, and more; see the SKA’s Software and Computing page for more information) is the big challenge that we want to address with these grants.
AstroCompute Grants
The next step is the target of the grants program that I am sharing with you today. The AWS Scientific Computing group, along with our friends at the SKA, want to make sure that the rest of the world is ready to process this astronomical (sorry) amount of data when it starts to become available in 2020.
To this end, we are providing grants (AWS credits) and up to one petabyte of storage for an AWS Public Data Set. The data set will be initially provided by several of the SKA’s precursor telescopes including CSIRO’s ASKAP, the MWA in Australia, and KAT-7 (pathfinder to the SKA precursor telescope Meerkat) in South Africa. These telescopes have already seen first light and are now producing data. Over time, the data set will grow to the full petabyte using data provided by the other SKA partners. The grants are open to anyone who is making use of radio astronomical telescopes or radio astronomical data resources around the world.
The grants will be administered by the SKA. They will be looking for innovating, cloud-based algorithms and tools that will be able to handle and process this never ending data stream. You can also read their post on Seeing Stars Through the Cloud.
If you meet the basic qualifications listed above, have an interest in working on this problem, and would like to apply for a grant, please visit the Call for Proposals page.
— Jeff;
Next Generation Genomics With AWS
My colleague Matt Wood wrote a great guest post to announce new support for one of our genomics partners.
— Jeff;
I am happy to announce that AWS will be supporting the work of our partner, Seven Bridges Genomics, who has been selected as one of the National Cancer Institute (NCI) Cancer Genomics Cloud Pilots. The cloud has become the new normal for genomics workloads, and AWS has been actively involved since the earliest days, from being the first cloud vendor to host the 1000 Genomes Project, to newer projects like designing synthetic microbes, and development of novel genomics algorithms that work at population scale. The NCI Cancer Genomics Cloud Pilots are focused on how the cloud has the potential to be a game changer in terms of scientific discovery and innovation in the diagnosis and treatment of cancer.
The NCI Cancer Genomics Cloud Pilots will help address a problem in cancer genomics that is all too familiar to the wider genomics community: data portability. Today’s typical research workflow involves downloading large data sets, (such as the previously mentioned 1000 Genomes Project or The Cancer Genome Atlas (TCGA)) to on-premises hardware, and running the analysis locally. Genomic datasets are growing at an exponential rate and becoming more complex as phenotype-genotype discoveries are made, making the current workflow slow and cumbersome for researchers. This data is difficult to maintain locally and share between organizations. As a result, genomic research and collaborations have become limited by the available IT infrastructure at any given institution.
The NCI Cancer Genomics Cloud Pilots will take the natural step to solve this problem, by bringing the computation to where the data is, rather than the other way around. The goal of the NCI Cancer Genomics Cloud Pilots are to create cloud-hosted repositories for cancer genome data that reside alongside the tools, algorithms, and data analysis pipelines needed to make use of the data. These Pilots will provide ways to provision computational resources within the cloud so that researchers can analyze the data in place. By collocating data in the cloud with the necessary interface, algorithms, and self-provisioned resources, these Pilots will remove barriers to entry, allowing researchers to more easily participate in cancer research and accelerating the pace of discovery. This means more life-saving discoveries such as better ways to diagnose stomach cancer, or the identification of novel mutations in lung cancer that allow for new drug targets.
The Pilots will also allow cancer researchers to provision compute clusters that change as their research needs change. They will have the necessary infrastructure to support their research when they need it, rather than make a guess at the resources that they will need in the future every time grant writing season starts. They will also be able to ask many more novel questions of the data, now that they are no longer constrained by a static set of computational resources.
Finally, the NCI Cancer Genomics Pilots will help researchers collaborate. When data sets are publicly shared, it becomes simple to exchange and share all the tools necessary to reproduce and expand upon another lab’s work. Other researchers will then be able to leverage that software within the community, or perhaps even in an unrelated field of study, resulting in even more ideas be generated.
Since 2009, Seven Bridges Genomics has developed a platform to allow biomedical researchers to leverage AWS’s cloud infrastructure to focus on their science rather than managing computational resources for storage and execution. Additionally, Seven Bridges has developed security measures to ensure compliance with Health Insurance Portability and Accountability Act (HIPAA) for all data stored in the cloud. For the NCI Cancer Genomics Cloud Pilots, the team will adapt the platform to meet the specific needs of the cancer research community as the develop over the course of the Pilot. If you are interested in following the work being done by Seven Bridges Genomics or giving feedback as their work on the NCI Cancer Genomics Cloud Pilots progresses, you can do so here.
We look forward to the journey ahead with Seven Bridges Genomics. You can learn more about AWS and Genomics here.
— Matt Wood, General Manager, Data Science
Big Data Update – New Blog and New Web-Based Training
The topic of big data comes up almost every time I meet with current or potential AWS customers. They want to store, process, and extract meaning from data sets that seemingly grow in size with every meeting.
In order to help our customers to understand the full spectrum of AWS resources that are available for use on their big data problems, we are introducing two new resources — a new AWS Big Data Blog and web-based training on Big Data Technology Fundamentals.
AWS Big Data Blog
The AWS Big Data Blog is a way for data scientists and developers to learn big data best practices, discover which managed AWS Big Data services are the best fit for their use case, and help them get started on AWS big data services. Our goal is to make this the hub for developers to discover new ways to collect, store, clean, process, and visualize data at any scale.
Readers will find short tutorials with code samples, case studies that demonstrate the unique benefits of doing big data on AWS, new feature announcements, partner- and customer-generated demos and tutorials, and tips and best practices for using AWS big data services.
The first two posts on the blog show you how to Build a Recommender with Apache Mahout on Amazon Elastic MapReduce and how to Power Gaming Applications with Amazon DynamoDB.
Big Data Training
If you are looking for a structured way to learn more about the tools, techniques, and options available to you as you learn more about big data, our new web-based Big Data Technology Fundamentals course should be of interest to you.
You should plan to spend about three hours going through this course. You will first learn how to identify common tools and technologies that can be used to create big data solutions. Then you will gain an understanding of the MapReduce framework, including the map, shuffle and sort, and reduce components. Finally, you will learn how to use the Pig and Hive programming frameworks to analyze and query large amounts of data.
You will need a working knowledge of programming in Java, C#, or a similar language in order to fully benefit from this training course.
The web-based course is offered at no charge, and can be used on its own or to prepare for our instructor-led Big Data on AWS course.
— Jeff;
Process Earth Science Data on AWS With NASA / NEX Public Data Sets
We have been working with the NASA Earth Exchange (NEX) team to make it easier and more efficient for researchers to access and process earth science data. The goal is to make a number of important data sets accessible to a wider audience of full-time researchers, students, and citizen scientists. This important new project is called OpenNEX.
Up until now, it has been logistically difficult for researchers to gain easy access to this data due to its dynamic nature and immense size (tens of terabytes). Limitations on download bandwidth, local storage, and on-premises processing power made in-house processing impractical.
Today we are publishing an initial collection of datasets available (over 20 TB), along with Amazon Machine Images (AMIs), and tutorials. NASA is also planning to host a series of virtual workshops for those interested in learning more about the datasets and how to process them on AWS.
The datasets are stored in Amazon S3 are can be found at s3://nasanex. Let’s take a look at each one…
Data for Climate Assessment
Formally known as the NASA Earth Exchange Downscaled Climate Projections, this dataset is described as follows:
The NASA Earth Exchange (NEX) Downscaled Climate Projections (NEX-DCP30) dataset is comprised of downscaled climate scenarios for the conterminous United States that are derived from the General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 5 (CMIP5) and across the four greenhouse gas emissions scenarios known as Representative Concentration Pathways (RCPs) developed for the Fifth Assessment Report of the Intergovernmental Panel on Climate Change (IPCC AR5). The dataset includes downscaled projections from 33 models, as well as ensemble statistics calculated for each RCP from all model runs available. The purpose of these datasets is to provide a set of high resolution, bias-corrected climate change projections that can be used to evaluate climate change impacts on processes that are sensitive to finer-scale climate gradients and the effects of local topography on climate conditions. Each of the climate projections includes monthly averaged maximum temperature, minimum temperature, and precipitation for the periods from 1950 through 2005 (Retrospective Run) and from 2006 to 2099 (Prospective Run).
You can access this dataset at s3://nasanex/NEX-DCP30. Consult the detail page and the tech note to learn more about the provenance, format, structure, and attribution requirements.
Landsat Global Land Survey
Landsat has been acquiring space-based moderate-resolution land remote sensing data for the past four decades. This is a unique resource for those who work in agriculture, geology, forestry, regional planning, education, mapping, and global change research. The Landsat images are also invaluable for emergency response and disaster relief.
Here’s the official description:
In the past, the U.S. Geological Survey (USGS) and NASA collaborated on the creation of four global land data sets from Landsat images: one from the 1970s, and one each from circa 1990, 2000, and 2005. Each of these global data sets was created from the primary Landsat sensor in use at the time: the Multispectral Scanner (MSS) in the 1970s, the Thematic Mapper (TM) in 1990, Enhanced Thematic Mapper Plus (ETM+) in 2000, and a combination of TM and ETM+ in 2005.
You can access this dataset at s3://nasanex/Landsat. Consult the detail page and the project description to learn more. You can also use the Landsat tools to access and view the datasets.
MODIS Vegetation Indices
Formally known as MOD13Q1 (Vegetation Indices 16-Day L3 Global 250m), this dataset is described as follows:
Due to their simplicity, ease of application, and widespread familiarity, vegetation indices have a wide range of usage within the user community. Some of the more common applications may include global biogeochemical and hydrologic modeling, agricultural monitoring and forecasting, land-use planning, land cover characterization, and land cover change detection. Global MODIS vegetation indices are designed to provide consistent spatial and temporal comparisons of vegetation conditions. Blue, red, and near-infrared reflectances, centered at 469-nanometers, 645-nanometers, and 858-nanometers, respectively, are used to determine the MODIS daily vegetation indices. The MODIS Normalized Difference Vegetation Index (NDVI) complements NOAA’s Advanced Very High Resolution Radiometer (AVHRR) NDVI products and provides continuity for time series historical applications. MODIS also includes a new Enhanced Vegetation Index (EVI) that minimizes canopy background variations and maintains sensitivity over dense vegetation conditions. The EVI also uses the blue band to remove residual atmosphere contamination caused by smoke and sub-pixel thin cloud clouds. The MODIS NDVI and EVI products are computed from atmospherically corrected bi-directional surface reflectances that have been masked for water, clouds, heavy aerosols, and cloud shadows. Global MOD13Q1 data are available every 16 days at 250-meter spatial resolution as a gridded level-3 product in the Sinusoidal projection. Lacking a 250m blue band, the EVI algorithm uses the 500m blue band to correct for residual atmospheric effects, with negligible spatial artifacts. Vegetation indices are used for global monitoring of vegetation conditions and are used in products displaying land cover and land cover changes. These data may be used as input for modeling global biogeochemical and hydrologic processes and global and regional climate. These data also may be used for characterizing land surface biophysical properties and processes, including primary production and land cover conversion.
You can access this dataset at s3://nasanex/MODIS. Consult the detail page and the data description to learn more. The MODIS tools may also prove to be helpful.
Webification Data Access Framework
In conjunction with today’s AWS/NASA hackathon, NASA has published an open source tool called Webification (w10n for short). This open source tool simplifies access to data sets such as those described above. All data is accessed via URL and returned in JSON or binary format.
The Webification tool is available as a web service, an EC2 AMI (ami-fc0f97cc in US West (Oregon)), and in source code form. Click through the interactive Webification tutorial to get started.
There’s also a visual aspect to the Webification tool. After you extract some data and convert it to JSON, you can create interactive, embedded visualizations that look like this:
AWS CLI Access
The AWS Command Line Interface (CLI) can also be used to access the datasets. Install and configure it, and then review the datasets like this:
$ aws s3 ls s3://nasanex Bucket: nasanex Prefix: LastWriteTime Length Name ------------- ------ ---- PRE Landsat/ PRE MODIS/ PRE NEX-DCP30/
Issue a series of ls commands to explore the bucket, and then download the file or files of interest:
$ aws s3 cp s3://nasanex/NEX-DCP30/NEX-quartile/rcp26/mon/atmos/tasmax/r1i1p1/v1.0/CONUS/tasmax_quartile75_amon_rcp26_CONUS_209601-209912.nc . download: s3://nasanex/NEX-DCP30/NEX-quartile/rcp26/mon/atmos/tasmax/r1i1p1/v1.0/CONUS/tasmax_quartile75_amon_rcp26_CONUS_209601-209912.nc to tasmax_quartile75_amon_rcp26_CONUS_209601-209912.nc $ ls -l tasmax_quartile75_amon_rcp26_CONUS_209601-209912.nc -rw-rw-r-- 1 jbarr jbarr 1088838203 Sep 29 08:03 tasmax_quartile75_amon_rcp26_CONUS_209601-209912.nc
Learn More
You can learn more about the OpenNEX project at the NASA OpenNEX page and on the AWS Public Data Sets page.
— Jeff;
The AWS Report – Matt Wood Discusses Big Data and re:Invent
In this episode of The AWS Report, I spoke with AWS Evangelist Matt Wood to learn about his track on Big Data and Analytics at AWS re:Invent. The track sounds awesome and I hope to be able to attend some of the sessions:
See you at re:Invent!
— Jeff;
SAP HANA One – Now Available for Production Use on AWS
Earlier this year I briefly mentioned SAP HANA and the fact that it was available for developer use on AWS.
Today, SAP announced HANA One, a deployment option for HANA that is certified for production use on AWS available now in the AWS Marketplace. You can run this powerful, in-memory database on EC2 for just $0.99 per hour.
Because you can now launch HANA in the cloud, you don’t need to spend time negotiating an enterprise agreement, and you don’t have to buy a big server. If you are running your startup from a cafe or commanding your enterprise from a glass tower, you get the same deal. No long-term commitment and easy access to HANA, on an hourly, pay-as-you-go basis, charged through your AWS account.
What’s HANA?
SAP HANA is an in-memory data platform well suited for performing real-time analytics, and developing and deploying real-time applications.
I spent some time watching the videos on the Experience HANA site as I was getting ready to write this post. SAP founder Hasso Plattner described the process that led to the creation of HANA, starting with a decision to build a new enterprise database in December of 2006. He explained that he wanted to capitalize on two industry trends — the availability of multi-core CPUs and the growth in the amount of RAM per system. Along with this, he wanted to exploit parallelism within the confines of a single application. Here’s what they came up with:
Putting it all together, SAP HANA runs entirely in memory, eschewing spinning disk entirely except for backup. Traditional disk-based data management solutions are optimized for transactional or analytic processing, but not both. Transactional processing is oriented around and optimized for row-base operations: inserts, updates, and deletes. In contrast, analytic processing is tuned for complex queries, often involving subsets of the columns in a particular table (hence the rise of column-oriented databases). All of this specialization and optimization is needed due to the fact that accessing data stored on a disk is 10,000 to 1,000,000 times slower than accessing data stored in memory. In addition to this bottleneck, disk-based systems are unable to take full advantage of multi-core CPUs.
At the base, SAP HANA is a complete, ACID-compliant relational database with support for most of SQL-92. At the top, you’ll find an analytical interface using Multi-Dimensional Expressions (MDX) and support for SAP BusinessObjects. Between the two is a parallel data flow computing engine designed to scale across cores. HANA also includes a Business Function Library, a Predictive Analysis Library, and the “L” imperative language.
So, what is HANA good for? Great question! Here are some applications:
Real-time analytics such as data warehousing, predictive analysis on Big Data, and operational (sales, finance, or shipping) reporting.
Real-time applications such as core process (e.g. ERP) acceleration, planning and optimization, and sense and response (smart meters, point of sale, and the like).
As an example of what can be done, SAP Expense Insight uses HANA and it is also available in the AWS Marketplace. It offers budget visibility to department managers in real-time, across any time horizon.
The folks at Taulia are building a dynamic discounting platform around HANA One. They’re already using AWS to streamline their deployment and operations; HANA One will allow them to make their platform even more responsive.
This is an enterprise-class product (but one that’s accessible to everyone) and I’ve barely scratched the surface. You can read this white paper to learn more (you may have to give the downloaded file a “.pdf” extension in order to open it).
Deploy HANA Now
As I mentioned earlier, SAP has certified HANA for production use on AWS. You can launch it today and you can get started now.
You don’t have to spend a lot of money. You don’t need to buy and install high-end hardware in you data center and you don’t need to license HANA. Instead, you can launch HANA from the AWS Marketplace and pay for the hardware and the software on an hourly, pay-as-you-go basis.
You’ll pay $0.99 per hour to run HANA One on AWS, plus another $2.50 per hour for an EC2 Cluster Compute Eight Extra Large instance with 60.5 GB of RAM and dual Intel Xeon E5 processors, bringing the total software and hardware cost to just $3.49 per hour, plus standard AWS fees for EBS and data transfer.
To get started, visit the SAP HANA page in the AWS Marketplace.
— Jeff;
Scaling Science: 1 Million Compute Hours in 1 week
For many scientists, the computer has become as important as the test tube, the centrifuge or the grad student in delivering ground breaking research. Whether screening for active cancer treatments or colliding atoms, the availability of compute cycles can significantly affect the time it takes for scientists to crunch their numbers. Indeed, compute resources are often so constrained that researchers often have to scale back the scope of their work to fit the capacity available.
Not so with Amazon EC2, where the general purpose, utility computing model is a perfect fit for scientific workloads of any scale. Researchers (and their grad students), can access the computational resources they need to deliver on their scientific vision, while staying focused on their analysis and results.
Scaling up at the Morgridge Instutute
Victor Ruotti faced this exact problem. His team at the Morgridge Institute at the University of Wisconsin-Madison are looking at the genes expressed as template cells, stem cells, start to take on the various special functions our tissues need, such as absorbing nutrients or conducting nervous impulses. Impressive and important work, which has large computational requirements: millions of RNA sequence reads and a data footprint of 78 TB.
Victor’s research was selected as the winner of Cycle Computing’s inaugural Big Science Challenge, and using’s Cycle’s software ran through the 15,376 alignment runs on Amazon EC2, clocking up over a million compute hours in a week, for just $116 an hour.
A Century of Compute
Over 1,000,000 compute hours, 115 years of work for a single processor, were used to build the genetic map the team needed to quickly identify which regions of the genome are important for establishing cell types which have clinical importance. The entire analysis started running on Spot instances in just 20 minutes, on high memory instance types (the M2 class), meaning that the team could use Cycle Server to stretch their budget further and build an extremely high resolution genetic map. The spot price was typically 12 times lower than the equivalent on-demand price, and their cluster ran across an average of 5000 instances (8000 at peak), for a total cost of $19,555. That’s less than the price of 20 lab pipettes.
Cycle Computing on the AWS Report
Our very own Jeff Barr was lucky enough to spend a few minutes chatting with Cycle Computing CEO, Jason Stowe for the AWS Report. Here is the episode they recorded:
Cycle also have a blog post with some more information on this, and the 2012 Big Science Challenge.
We’re very happy to see the utility computing platform of AWS be used for such ground breaking work. If you’re working with data and would like to discuss how to get up and running at this, or any other scale, I do hope you’ll get in touch.
Upcoming Webinar
If you would like to know more I’ll be hosting a webinar on big data and HPC on the 16th of October. We’ll discuss some customer success stories and common best practices for using tools such as Elastic MapReduce, DynamoDB and the broad range of services on the AWS Marketplace to accelerate your own applications and analytics.
Registration is free. See you there.
~ Matt

