AWS Official Blog

  • Amazon EMR Release 4.1.0 โ€“ Spark 1.5.0, Hue 3.7.1, HDFS Encryption, Presto, Oozie, Zeppelin, Improved Resizing

    by Jeff Barr | on | in Amazon EMR |

    My colleagues Jon Fritz and Abhishek Sinha are both Senior Product Managers on the EMR team. They wrote the guest post below to introduce you to the newest release of EMR and to tell you about new EMR cluster resizing functionality.

    โ€” Jeff;


    Amazon EMR is a managed service that simplifies running and managing distributed data processing frameworks, such as Apache Hadoop and Apache Spark.

    Today we are announcing Amazon EMR release 4.1.0, which includes support for Spark 1.5.0, Hue 3.7.1 and HDFS transparent encryption with Hadoop KMS. We are also introducing an intelligent resize feature that allows you to reduce the number of nodes in your cluster with minimal impact to running jobs. Finally, we are also announcing the availability of Presto 0.119, Zeppelin 0.6 (Snapshot) and Oozie 4.0.1 as Sandbox Applications. The EMR Sandbox gives you early access to applications which are still in development for a full General Availability (GA) release.

    EMR release 4.1.0 is our first follow-up release to 4.0.0, which brought many new platform improvements around configuration of applications, a new packaging system, standard ports and paths for Hadoop ecosystem applications, and a Quick Create option for clusters in the AWS Management Console.

    New Applications and Components in the 4.x Release Series
    Amazon EMR provides an easy way to install and configure distributed big data applications in the Hadoop and Spark ecosystems on your cluster when creating clusters from the EMR console, AWS CLI, or using a SDK with the EMR API. In release 4.1.0, we have added support for several new applications:

    • Spark 1.5.0 โ€“ We included Spark 1.4.1 on EMR release 4.0.0, and we have upgraded the version of Spark to 1.5.0 in this EMR release. Spark 1.5.0 includes a variety of new features and bug fixes, including additional functions for Spark SQL/Dataframes, new algorithms in MLlib, improvements in the Python API for Spark Streaming, support for Parquet 1.7, and preferred locations for dynamically allocated executors. To learn more about Spark in Amazon EMR, click here.
    • HUE 3.7.1 โ€“ Hadoop User Experience (HUE) is an open source user interface which allows users to more easily develop and run queries and workflows for Hadoop ecosystem applications, view tables in the Hive Metastore, and browse files in Amazon S3 and on-cluster HDFS. Multiple users can login to HUE on an Amazon EMR cluster to query data in Amazon S3 or HDFS using Apache Hive and Pig, create workflows using Oozie, develop and save queries for later use, and visualize query results in the UI. For more information about how to connect to the HUE UI on your cluster, click here.
    • Hadoop KMS for HDFS Transparent Encryption โ€“ The Hadoop Key Management Server (KMS) can supply keys for HDFS Transparent Encryption, and it is installed on the master node of your EMR cluster with HDFS. You can also use a key vendor external to your EMR cluster which utilizes the Hadoop KeyProvider API. Encryption in HDFS is transparent to applications reading from and writing to HDFS, and data is encrypted in in-transit in HDFS because encryption and decryption activities are carried out in the client. Amazon EMR has also included an easy configuration option to programmatically create encrypted HDFS directories when launching clusters. To learn more about using Hadoop KMS with HDFS Transparent Encryption, click here.

    Introducing the EMR Sandbox
    With the EMR Sandbox, you now have early access to new software for your EMR cluster while those applications are still in development for a full General Availability (GA) release. Previously, bootstrap actions were the only mechanism to install applications not fully supported on EMR. However, you would need to specify a bootstrap action script, the installation was not tightly coupled to an EMR release, and configuration settings were harder to maintain. Instead, applications in the EMR Sandbox are certified to install correctly, configured using a configuration object, and specified directly from the EMR console, CLI, or EMR API using the application name (ApplicationName-Sandbox). Release 4.1.0 has three EMR Sandbox applications:

    • Presto 0.119 โ€“ Presto is an open-source, distributed SQL query engine designed to query large data sets in one or more heterogeneous data sources, including Amazon S3. Presto is optimized for ad-hoc analysis at interactive speed and supports standard ANSI SQL, including complex queries, aggregations, joins, and window functions. Presto does not use Hadoop MapReduce; instead, it uses a query execution mechanism that processes data in memory and pipelines it across the network between stages. You can interact with Presto using the on-cluster Presto CLI or connect with a supported UI like Airpal, a web-based query execution tool which was open sourced by Airbnb. Airpal has several interesting features such as syntax highlighting, results exported to a CSV for download, query history, saved queries, table finder to search for appropriate tables, and a table explorer to visualize schema of a table and sample the first 1000 rows. To learn more about using Airpal with Presto on Amazon EMR, read the new post, Analyze Data with Presto and Airpal on Amazon EMR on the AWS Big Data Blog. To learn more about Presto on EMR, click here.
    • Zeppelin 0.6 (Snapshot) โ€“ Zeppelin is an open source GUI which creates interactive and collaborative notebooks for data exploration using Spark. You can use Scala, Python, SQL (using Spark SQL), or HiveQL to manipulate data and quickly visualize results. Zeppelin notebooks can be shared among several users, and visualizations can be published to external dashboards. When executing code or queries in a notebook, you can enable dynamic allocation of Spark executors to programmatically assign resources or change Spark configuration settings (and restart the interpreter) in the Interpreter menu.
    • Oozie 4.0.1 โ€“ Oozie is a workflow scheduler for Hadoop, where you can create Directed Acyclic Graphs (DAGs) of actions. Also, you can easily trigger your Hadoop workflows by actions or time.

    Example Customer Use Cases for Presto on Amazon EMR
    Even before Presto was supported as a Sandbox Application, many AWS customers have been using Presto on Amazon EMR, especially for interactive ad hoc queries on large scale data sets in Amazon S3. Here are a few examples:

    • Cogo Labs, a startup incubator, operates a platform for marketing analytics and business intelligence. Presto running on Amazon EMR allows any of their 100+ developers and analysts to run SQL queries on over 500 TB of data stored in Amazon S3 for data-exploration, ad-hoc analysis, and reporting.
    • Netflix has chosen Presto as their interactive, ANSI-SQL compliant query engine for big data, as Presto scales well, is open source, and integrates with the Hive Metastore and Amazon S3 (the backbone of Netflixโ€™s Big Data Warehouse environment.) Netflix runs Presto on persistent EMR clusters to quickly and flexibly query across their ~25PB S3 data store. Netflix is an active contributor to Presto, and Amazon EMR provides Netflix with the flexibility to run their own build of Presto on Amazon EMR clusters. On average, Netflix runs ~3500 queries per day on their Presto clusters. Learn more about Netflixโ€™s Presto deployment.
    • Jampp is a mobile application marketing platform, and they use advertising retargeting techniques to drive engaged users to new applications. Jampp currently uses Presto on EMR to process 40 TB of data each day.
    • Kanmu is Japanese startup in the financial services industry and provides offers based on consumersโ€™ credit card usage. Kanmu migrated from Hive to using Presto on Amazon EMR because of Prestoโ€™s ability to run exploratory and iterative analytics at interactive speeds, good performance with Amazon S3, and scalability to query large data sets.
    • OpenSpan provides automation and intelligence solutions that help bridge people, processes and technology to gain insight into employee productivity, simplify transactions, and engage employees and customers. OpenSpan migrated from HBase to Presto on Amazon EMR with Amazon S3 as a data layer. OpenSpan chose Presto because of its ANSI SQL interface and ability to query data in real-time directly from Amazon S3, which allows them to quickly explore vast amounts of data and rapidly iterate on upcoming data products.

    Intelligent Resize Feature Set
    In release 4.1.0, we have added an Intelligent Resize feature set so you can now shrink your EMR cluster with minimal impact to running jobs. Additionally, when adding instances to your cluster, EMR can now start utilizing provisioned capacity as soon it becomes available. Previously, EMR would need the entire requested capacity to become available before allowing YARN to send tasks to those nodes. Also, you can now issue a resize request to EMR while a current resize request is being executed (to change the target size of your cluster), or stop a resize operation.

    When decreasing the size of your cluster, EMR will programmatically select instances which are not running tasks, or if all instances in the cluster are being utilized, EMR will wait for tasks to complete on a given instance before removing it from the cluster. The default wait time is 1 hour, and this value can be changed. You can also specify a timeout value in seconds by changing the yarn.resourcemanager.decommissioning parameter in /home/hadoop/conf/yarn-site.xml or /etc/hadoop/conf/yarn-site.xml file. EMR will dynamically update the new setting and a resource manager restart is not required. You can set this to arbitrarily large number to ensure that no tasks are killed while shrinking the cluster.

    Additionally, Amazon EMR now has support for removing instances in the core group, which store data as a part of HDFS along with running YARN components. When shrinking the number of instances in a clusterโ€™s core group, EMR will gracefully decommission HDFS daemons on the instances. During the decommissioning process, HDFS replicates the blocks on that instance to other active instances to reach the desired replication factor in the cluster (EMR sets the default replication factor to 1 for 1-3 core nodes, the value to 2 for 4-9 core nodes, and the value to 3 for 10+ core nodes). To avoid data loss, EMR will not allow shrinking your core group below the storage required by HDFS to store data on the cluster, and will ensure that the cluster has enough free capacity to successfully replicate blocks from the decommissioned instance to the remaining instances. If the requested instance count is too low to fit existing HDFS data, only a partial number of instances will be decommissioned.

    We recommend minimizing HDFS heavy writes before removing nodes from your core group. HDFS replication can slow down due to under-construction blocks and inconsistent replica blocks, which will decrease the performance of the overall resize operation. To learn more about resizing your EMR clusters, click here.

    Launch an Amazon EMR Cluster With 4.1.0 Today
    To create an EMR cluster with 4.1.0, select release 4.1.0 on the Create Cluster page in the AWS Management Console, or use the release label โ€œemr-4.1.0โ€ when creating your cluster from the AWS CLI or using a SDK with the EMR API.

    โ€” Jon Fritz and Abhishek Sinha

  • New AWS Security Courses (Fundamentals & Operations)

    by Jeff Barr | on | in Training and Certification |

    Itโ€™s probably no surprise that information security is one of todayโ€™s most sought after IT specialties. Itโ€™s also deeply important to our customers and any company considering moving to the cloud.

    So, today weโ€™re launching a new AWS Training curriculum focused on security. The curriculumโ€™s two new classes are designed to help you meet your cloud security objectives under the AWS Shared Responsibility Model, by showing you how to create more secure AWS architectures and solutions and address key compliance requirements.

    Hereโ€™s a closer look at whatโ€™s new:

    • AWS Security Fundamentals โ€“ This free 3-hour online class is designed to introduce fundamental cloud computing and AWS security concepts, including AWS access control and management, governance, logging, and encryption methods. The class, aimed primarily at security professionals with little or no working knowledge of AWS, also addresses security-related compliance protocols, risk management strategies, and procedures for auditing AWS security infrastructure.
    • Security Operations on AWS โ€“ A 3-day technical deep dive on how to stay secure and compliant in the AWS cloud. This classroom-based course covers security features of key AWS services and AWS best practices for securing data and systems. Youโ€™ll learn about regulatory compliance standards and use cases for running regulated workloads on AWS. Hands-on practice with AWS security products and features will help you take your security operations to the next level.

    Visit AWS Training to learn more about the new security courses and find an instructor-led class near you.

    โ€” Jeff;

  • New AWS Digital Library for Big Data Solutions

    by Jeff Barr | on | in Big Data, Case Studies | | Comments

    My colleague Luis Daniel Soto has been working with AWS Community Hero Lynn Langit to create a comprehensive collection of resources for customers who are ready to run Big Data applications on AWS!

    Hereโ€™s what they have to sayโ€ฆ.

    โ€” Jeff;


    Today the AWS Marketplace is launching a new on-line video library designed to help our customers find AWS Marketplace vendor solutions, as well as accelerate and manage short and long-term data integration, business intelligence and advanced analytics projects for their AWS cloud and on-premises data.

    The AWS Marketplace Digital Library for Big Data provides business and technical content from AWS Marketplace technology vendors and case studies from customers who have built end-to-end Big Data solutions. The segments are hosted by cloud and Big Data architect Lynn Langit and organized around a common set of functionality to help organizations and individuals find the AWS Marketplace vendor solutions to address their particular needs.

    The library is hosted on a video webcasting platform which allows our customers to interact with AWS Marketplace partners, by asking questions as they watch the demos and interviews in split-screen mode. Hereโ€™s a sample:

    If you are an APN Partner and want to learn more or want to be part of the AWS Digital Library, visit the new Big Data Partner Solutions page.

    โ€” Luis and Lynn

  • New โ€“ Receive and Process Incoming Email with Amazon SES

    by Jeff Barr | on | in Simple Email Service | | Comments

    We launched the Amazon Simple Email Service (SES) way back in 2011, with a focus on deliverability โ€” getting mail through to the intended recipients. Today, the service is used by Amazon and our customers to send billions of transactional and marketing emails each year.

    Today we are launching a much-requested new feature for SES. You can now use SES to receive email messages for entire domains or for individual addresses within a domain. This will allow you to build scalable, highly automated systems that can programmatically send, receive, and process messages with minimal human intervention.

    You use sophisticated rule sets and IP address filters to control the destiny of each message. Messages that match a rule can be augmented with additional headers, stored in an S3 bucket, routed to an SNS topic, passed to a Lambda function, or bounced.

    Receiving and Processing Email
    In order to make use of this feature you will need to verify that you own the domain of interest. If you have already done this in order to use SES to send email, then you are already good to go.

    Now you need to route your incoming email to SES for processing. You have two options here. You can set the domainโ€™s MX (Mail Exchange) record to point to the SES SMTP endpoint in the region where you want to process incoming email. Or, you can configure your existing mail handling system to forward mail to the endpoint.

    The next step is to figure out what you want to do with the messages. To do this, you need to create some receipt rules. Rules are grouped into rule sets (order matters within the set) and can apply to multiple domains. Like most aspects of AWS, rules and rule sets are specific to a particular region. You can have one active rule set per AWS region; if you have no such set, then all incoming email will be rejected.

    Rules have the following attributes:

    • Enabled โ€“ A flag that enables or disables the rule.
    • Recipients โ€“ A list of email addresses and/or domains that the rule applies to. If this attribute is not supplied, the rule matches all addresses in the domain.
    • Scan โ€“ A flag to request spam and virus scans (default is true).
    • TLS โ€“ A flag to require that mail matching this rule is delivered over a connection that is encrypted with TLS.
    • Action List -An ordered list of actions to perform on messages that match the rule.

    When SES receives a message, it performs several checks before it accepts the message for further processing. Hereโ€™s what happens:

    • The source IP address is checked against an internal list maintained by SES, and rejected if so (this list can be overridden using an IP address filter that explicitly allows the IP address).
    • The source IP address is checked against your IP address filters, and rejected if so directed by the filter.
    • The message is checked to see if it matches any of the recipients specified in a rule, or if thereโ€™s a domain level match, and accepted if so.

    Messages that do not match a rule do not cost you anything. After a message has been accepted, SES will perform the actions associated with the matching rule.  The following actions are available:

    • Add a header to the message.
    • Store the message in a designated S3 bucket, with optional encryption using a key stored in AWS Key Management Service (KMS). The entire message (headers and body) must be no larger than 30 megabytes in size for this action to be effective.
    • Publish the message to a designated SNS topic. The entire message (headers and body) must be no larger than 150 kilobytes in size for this action to be effective.
    • Invoke a Lambda function. The invocation can be synchronous or asynchronous (the default).
    • Return a specified bounce message to the sender.
    • Stop processing the actions in the rule.

    The actions are run in the order specified by the rule. Lambda actions have access to the results of the spam and virus scans and can take action accordingly. If the Lambda function needs access to the body of the message, a preceding action in the rule must store the message in S3.

    A Quick Demo
    Hereโ€™s how I would create a rule that passes incoming email messages to a Lambda function (MyFunction) notifies an SNS topic (MyTopic), and then stores the messages in an S3 bucket (MyBucket) after encrypting them with a KMS key (aws/ses):

    I can see all of my rules at a glance:

    Hereโ€™s a Lambda function that will stop further processing if a message fails any of the spam or virus checks. In order for this function to perform as expected, it must be invoked in synchronous (RequestResponse) fashion.

    exports.handler = function(event, context) {
        console.log('Spam filter');
        
        var sesNotification = event.Records[0].ses;
        console.log("SES Notification:\n", JSON.stringify(sesNotification, null, 2));
     
        // Check if any spam check failed
        if (sesNotification.receipt.spfVerdict.status      === 'FAIL'
            || sesNotification.receipt.dkimVerdict.status  === 'FAIL'
            || sesNotification.receipt.spamVerdict.status  === 'FAIL'
            || sesNotification.receipt.virusVerdict.status === 'FAIL')
        {
            console.log('Dropping spam');
            // Stop processing rule set, dropping message
            context.succeed({'disposition':'STOP_RULE_SET'});
        }
        else
        {
            context.succeed();   
        }
    };

    To learn more about this feature, read Receiving Email in the Amazon SES Developer Guide.

    Pricing and Availability
    You will pay $0.10 for every 1000 emails that you receive. Messages that are 256 KB or larger are charged for the number of complete 256 KB chunks in the message, at the rate of $0.09 per 1000 chunks. A 768 KB message counts for 3 chunks. Youโ€™ll also pay for any S3, SNS, or Lambda resources that you consume. Refer to the Amazon SES Pricing page for more information.

    This new feature is available now and you can start using it today. Amazon SES is available in the US East (Northern Virginia), US West (Oregon), and Europe (Ireland) regions.

    โ€” Jeff;

  • Elastic Beanstalk Update โ€“ Support for Java and Go

    by Jeff Barr | on | in AWS Elastic Beanstalk |

    My colleague Abhishek Singh is the product manager for AWS Elastic Beanstalk. He wrote the following guest post in order to let you know that the service now supports Java JAR files and the Go programming language!

    โ€” Jeff;


    AWS Elastic Beanstalk already simplifies the process of deploying and scaling Java, .NET, PHP, Python, Ruby, Node.js, and Docker web applications and services on AWS. You simply upload your code and Elastic Beanstalk automatically handles the deployment, including capacity provisioning, load balancing, auto-scaling to application health monitoring. At the same time, you retain full control over the AWS resources powering your application and can access them at any time.

    Today we are making Elastic Beanstalk even more useful by adding support for Java and Go applications. In addition, the new platforms simplify the process of configuring the Nginx reverse proxy that runs on the web tier. You can now place an nginx.conf file in the .ebextensions/nginx folder to override the Nginx configuration. You can also place configuration files in the .ebextensions/nginx/conf.d folder in order to have them included in the Nginx configuration provided by the platform. For more information, see Configuring the Reverse Proxy.

    1. .ebextensions/nginx/nginx.conf โ€“ Overrides the Nginx configuration for the platform.
    2. .ebextensions/nginx/conf.d โ€“ Files are included in the Nginx configuration provided by the platform.

    New Support for Java
    You can now run any Java application, including those that use servers or frameworks such as Jetty or Play and are no longer restricted to using Tomcat as the application server for your Java applications.

    You can deploy your Java application to Elastic Beanstalk in the following ways:

    To get started, simply create a new Elastic Beanstalk environment and select the Java platform under the Preconfigured category. Both Java 7 and Java 8 are supported:

    New Support for Go
    Also, you can now run Go language applications on AWS Elastic Beanstalk. You can deploy your Go application to Elastic Beanstalk in the following ways:

    1. Upload an archive containing your applicationโ€™s source. AWS Elastic Beanstalk will automatically build and run your application (AWS Elastic Beanstalk assumes that the main function is in a file named application.go).
    2. Upload an archive containing your applicationโ€™s binary with a Procfile defining additional command line parameters required to run your application. See Application Process Configuration (Procfile) for details.
    3. Upload an archive containing your applicationโ€™s source, a Buildfile, and a Procfile. For details, see Building Applications On-Server (Buildfile).

    Like the Java platform, the Go platform also supports running multiple processes by defining them in a Procfile.

    To begin using the new platforms, log in to the AWS Elastic Beanstalk Management Console or use the EB CLI to create an environment running the appropriate platform.

    โ€” Abhishek Singh, Senior Product Manager, AWS Elastic Beanstalk

  • AWS Week in Review โ€“ September 21, 2015

    by Jeff Barr | on | in Week in Review |

    Letโ€™s take a quick look at what happened in AWS-land last week:

    Monday

    September 21

    Tuesday

    September 22

    Wednesday

    September 23

    Thursday

    September 24

    Friday

    September 25

    New & Notable Open Source

    • saws is a supercharged AWS Command Line Interface (CLI).
    • JAWS  is the server-less application framework.
    • aws-vault is a vault for securely storing and accessing AWS credentials in development environments.
    • iamy is an IAM import and export tool.
    • amazon-ecs-plugin is an EC2 Container Service plugin for Jenkins.
    • calypso is a set of tools for better Docker deployments to AWS Elastic Beanstalk.
    • s3-photo-archiver implements archival storage of photo and video media in S3.
    • shepherd is a framework for building APIs using AWS API Gateway and Lambda.
    • ish lets you SSH to an EC2 server based on name tag, AMI, autoscaling group, or instance ID.

    New SlideShare Presentations

    New Customer Success Stories

    • Canal+ -Mobile and personalized cable television offers.
    • Cinémur -Mobile app to rate and comment on movies.
    • Cydar -3D overlays for X-ray guided procedures.
    • Eyeota โ€“ Data collection and analysis for online publishers.
    • Healthcare.gov โ€“ Centers for Medicare and Medicaid Services.
    • Localytics -Marketing and analytics for major brands.
    • Peak โ€“ Brain training app.
    • Realeyes โ€“ Business intelligence via emotion analytics.

    New YouTube Videos

    New Marketplace Applications

    Upcoming Events

    Upcoming Events at the AWS Loft (San Francisco)

    Upcoming Events at the AWS Loft (New York)

    • September 28 โ€“ Onshape Users (6 PM โ€“ 8:30 PM).
    • October 1 โ€“ Behind the Scenes with LearnBop โ€“ Bulletproof Blue/Green Deployments โ€“ Myths, Pitfalls and Solutions (6:30 PM โ€“ 8 PM).
    • October 5 โ€“ A Brief (and Thorough) Introduction to Mindfulness (1 โ€“ 2:30 PM).
    • October 6 โ€“ AWS Pop-up Loft Trivia Night (6 โ€“ 8 PM).
    • October 7 โ€“ AWS re:Invent at the Loft โ€“ Keynote Live Stream (11:30 AM โ€“ 1:30 PM).
    • October 8 โ€“ AWS re:Invent at the Loft โ€“ Keynote Live Stream (12:00 PM โ€“ 1:30 PM).
    • October 8 โ€“ AWS re:Invent at the Loft โ€” re:Play Happy Hour! (7 โ€“ 9 PM).

    Upcoming Events at the AWS Loft (Berlin)

    • October 15 โ€“ An overview of Hadoop & Spark, using Amazon Elastic MapReduce (9 AM).
    • October 15 โ€“ Processing streams of data with Amazon Kinesis (and other tools) (10 AM).
    • October 15 โ€“ STUPS โ€“ A Cloud Infrastructure for Autonomous Teams (5 PM).
    • October 16 โ€“ Transparency and Audit on AWS (9 AM).
    • October 16 โ€“ Encryption Options on AWS (10 AM).
    • October 16 โ€“ Simple Security for Startups (6 PM).
    • October 19 โ€“ Introduction to AWS Directory Service, Amazon WorkSpaces, Amazon WorkDocs and Amazon WorkMail (9 AM).
    • October 19 โ€“ Amazon WorkSpaces: Advanced Topics and Deep Dive (10 AM).
    • October 19 โ€“ Building a global real-time discovery platform on AWS (6 PM).
    • October 20 โ€“ Scaling Your Web Applications with AWS Elastic Beanstalk (10 AM).

    Upcoming Events at the AWS Loft (London)

    • September 28 โ€“ Introduction to Funding and Pitching (5 PM).
    • September 29 โ€“ Masterclass Live: Amazon EC2 (10 AM).
    • September 29 โ€“ AWS CodeDeploy: Getting Started (1 PM).
    • September 29 โ€“ Amazon EC2 Container Service: Getting Started (3 PM).
    • September 29 โ€“ Cohesive Networks โ€“ Ensuring a secure foundation for your AWS Containers (4 PM).
    • September 29 โ€“ AWS for Startups (5 PM).
    • September 30 โ€“ IoT Lab Session (9 AM).
    • September 30 โ€“ Defining VPC Based Web Apps in AWS CloudFormation (10 AM).
    • September 30 โ€“ Startup Showcase โ€“ B2B (1 PM).
    • October 16 โ€“ HPC in the Cloud Workshop (2 โ€“ 4 PM).
    • October 22 โ€“ Working with Planetary-Scale Open Data Sets on AWS (2 โ€“ 4 PM).

    Help Wanted

    Stay tuned for next week! In the meantime, follow me on Twitter and subscribe to the RSS feed.

    โ€” Jeff;

  • Amazon Glacier Update โ€“ Third-Party SEC 17a-4(f) Assessment for Vault Lock

    by Jeff Barr | on | in Amazon Glacier |

    Amazon Glacier is designed to store any amount of archival or backup data with high durability.  Amazon Glacier is a very cost-effective solution (as low as $0.007 per gigabyte per month) for data that is infrequently accessed, and where a retrieval time of several hours is acceptable.

    Earlier this year we introduced a new Amazon Glacier compliance feature called Vault Lock (see my post, Create Write-Once-Read-Many Archive Storage with Amazon Glacier, to learn more).  As I wrote at the time, this feature allows you to lock your Amazon Glacier vaults with compliance controls that are designed (per SEC Rule 17a-4(f)) to help meet the requirement that โ€œelectronic records must be preserved exclusively in a non-rewritable and non-erasable format.โ€

    That announcement brought Amazon Glacier to the attention of AWS customers in the financial services industry.  Large banks, broker-dealers, and securities clearinghouses have all expressed interest in this important new feature.

    New Third-Party Assessment Report
    Today I am pleased to be able to announce that we have received a third-party assessment report that speaks to Amazon Glacierโ€™s ability to help meet the requirements of SEC 17a-4(f).

    This assessment is provided by Cohasset Associates, a highly respected consulting firm with more than 40 years of experience and knowledge related to the legal, technical, and operational issues associated with the records management practices of companies regulated by the US SEC (Securities and Exchange Commission) and the US CFTC (Commodity Futures Trading Commission).

    The full assessment (which is actually fairly interesting) provides a detailed look at the logic that Amazon Glacier uses to create immutable policies, along with a step-by-step examination and exposition of the controls that are used to protect Amazon Glacier vaults for compliance use cases once they have been locked (again, more information on this procedure can be found in the blog post that I referenced above).

    View the Amazon Glacier with Vault Lock Assessment to learn more. For information about other compliance features, visit the AWS Compliance Center.

    โ€” Jeff;

  • Amazon RDS Update โ€“ Oracle + Brazil + Larger Volumes + More

    by Jeff Barr | on | in Amazon RDS |

    I love to demo Amazon Relational Database Service (RDS) to live audiences! They always appreciate the fact that I can launch a MySQL, Oracle, SQL Server, PostgreSQL, or Amazon Aurora database instance with a couple of clicks.

    Today I would like to bring you up to date on a bunch of improvements that we have recently made to the service. I was not able to blog about these at launch time so this might not be news, but I did want to make sure that you didnโ€™t miss anything important. Hereโ€™s a quick summary of what I want to share with you:

    • The t2.large database instance type is now available.
    • Support for Oracle 12.1.0.2 and the latest patches is now available.
    • R3 and T2 database instances can now run Oracle.
    • The R3 database instances are now available in Brazil.
    • Database instances running MySQL, Oracle, SQL Server, and PostgreSQL can now be provisioned with even more storage (4 โ€“ 6 TB, depending on the database engine).
    • Tags on database instances are now copied to snapshots, and from there to instances restored from the snapshots.
    • You now have access to a license-included offering for SQL Server Enterprise Edition.

    Availability of t2.large Database Instances
    The T2 instances provide you with a baseline level of CPU performance and the ability to burst above the baseline. They are designed for workloads that do not need the entire CPU on a full or consistent basis, and are priced lower than comparable M3 DB instances.

    In addition to the existing instance types (db.t2.micro, db.t2.small, and db.t2.medium), you can now run all supported database engines on the new db.t2.large instance type. This instance type offers twice as much memory and 50% more CPU credits per hour than the db.t2.medium.  It is available in the US East (Northern Virginia), US West (Northern California), US West (Oregon), South America (Brazil), Europe (Ireland), Europe (Frankfurt), Asia Pacific (Tokyo), Asia Pacific (Sydney), Asia Pacific (Singapore), and China (Beijing) regions.

    The t2.large also supports encryption at rest. You can set this up on the Configure Advanced Settings page:

    Support for Oracle 12.1.0.2
    RDS for Oracle now supports version 12.1.0.2 of Oracle datatabase 12c. You can use the new In-Memory option to store a subset of your data in an in-memory column format that is optimized for performance. This is a great fit for the newly available R3 databases instances described in the next section.

    As part of this update, we also applied the April 2015 Oracle Patch Set Updates (PSU) for Oracle Database 11g and 12c and enabled access to the DBMS_REPAIR package. We also improved the integration with AWS CloudHSM; you can now access a single CloudHSM partition from multiple RDS accounts and you can store TDE master keys for multiple RDS Oracle databases on a single CloudHSM partition.

    You now have access to the following versions of Oracle through RDS:

    • 11.2.0.4.v4
    • 12.1.0.2.v1
    • 12.1.0.1.v2

    Oracle on R3 and T2 Database Instances
    The R3 instances are optimized for memory-intensive applications and have the lower cost per GiB of RAM of any DB instance. The instances deliver high sustained memory bandwidth and offer lower network latency, all at prices that are up to 28% lower than comparable M2 DB instances.

    You can now run Oracle Database on the R3 and T2 instances:

    R3 in Brazil
    The R3 database instances are now available in the South America (Brazil) Region, and can be used with the MySQL, Oracle, SQL Server, and PostgreSQL database engines.

    Provision Even More Storage
    Earlier this year we increased the amount of storage that you can provision when you use Provisioned IOPS or General Purpose (SSD) storage for an RDS database instance. Here are the new limits:

    • MySQL, PostgreSQL, and Oracle database instances can now be provisioned with up to 6 TB of storage.
    • SQL Server database instances can now be provisioned with up to 4 TB of storage and up to 20,000 IOPS (double the former limit).

    Instance Tags to Snapshots, and Back
    If you add tags to your database instances, create snapshots of those instances, and then use the snapshots to create fresh instances, the tags now appear on the new instances.

    SQL Server Enterprise, License Included
    You can now run SQL Server Enterprise Edition as a License Included offering on RDS. In other words, you do not need to purchase a separate license for the product; the pricing includes the software license, the underlying hardware resources, and the RDS management capabilities.

    Available Now
    These options are available now (some of them have been around for a month or two) and you can start using them today!

    โ€” Jeff;

     

  • In-Country Storage of Personal Data

    by Jeff Barr | on | in Security |

    My colleague Denis Batalov works out of the AWS Office in Luxembourg.  As a Solutions Architect, he is often asked about the in-country storage requirement that some countries impose on certain types of data. Although this requirement applies to a relatively small number of workloads, I am still happy that he took the time to write the guest post below to share some of his knowledge.

    โ€” Jeff;


    AWS customers sometimes offer their services in countries where local requirements necessitate storage and processing of certain sensitive data to take place within the applicable country, that is, in a datacenter physically located in the respective country. Examples of such sensitive data include financial transactions and personal data (also referred to in some countries as Personally Identifiable Information, or PII).  Depending on the specific storage and processing requirements, one answer might be to utilize hybrid architectures where the component of the system that is responsible for collecting, storing and processing the sensitive data is placed in-country, while the remaining system resides in AWS. More information about hybrid architectures in general can be found on the Hybrid Architectures with AWS page.

    The reference architecture diagram included below shows an example of a hypothetical web application hosted on AWS that collects personal data as part of its operation.  Since the collection of personal data may be required to occur in-country, the widget or form that is used to collect or display personal data (shown in red) is generated by a web server located in-country, while the rest of the web site (shown in green) is generated by the usual web server located in AWS. This way the authoritative copy of the personal data resides in-country and all updates to the data are also recorded in-country. Note that the data that is not required to be stored in-country can continue to be stored in the main database (or databases) residing in AWS.

    This architecture still provides customers with the most important benefits of the cloud: it is flexible, scalable, and cost-effective.

    There may be situations where a copy of personal data needs to be transferred across a national border, e.g. in order to fulfill contractual obligations, such as transferring the name, billing address and payment method when a cross-border purchase is transacted. Where permitted by local legislation, a replica of the data (either complete or partial) can be transferred across the border via a secure channel.  Data can be securely transferred over public internet with the use of TLS, or using a VPN connection established between the Virtual Private Gateway of the VPC and the Customer Gateway residing in-country.  Additionally, customers may establish private connectivity between AWS and an in-country datacenter by using AWS Direct Connect, which in many cases can reduce network costs, increase bandwidth throughput, and provide a more consistent network experience compared to Internet-based connections.

    Alternatively, it may be possible to achieve certain processing outcomes in the AWS cloud while employing data anonymization. This is a type of information sanitization whose intent is privacy protection, commonly associated with highly sensitive personal information. It is the process of either encrypting, tokenizing, or removing personally identifiable information from data sets, so that the people whom the data describe remain anonymous in a particular context. Upon return of the processed dataset from the AWS cloud it could be integrated in to in-country databases to give it personal context again.

    โ€” Denis

    PS โ€“ Customers should, of course, seek advice from professionals who are familiar with details of the country-specific legislation to ensure compliance with any applicable local laws, as this example architecture is shown here for illustrative purposes only!

  • Now Available โ€“ Amazon Linux AMI 2015.09

    by Jeff Barr | on | in Amazon EC2, Amazon Linux AMI | | Comments

    My colleague Max Spevack runs the team that produces the Amazon Linux AMI. He wrote the guest post below to announce the newest release!

    โ€” Jeff;


    The Amazon Linux AMI is a supported and maintained Linux image for use on Amazon EC2.

    We offer new major versions of the Amazon Linux AMI after a public testing phase that includes one or more Release Candidates. The Release Candidates are announced in the EC2 forum and we welcome feedback on them.

    Launching 2015.09 Today
    Today we announce the 2015.09 Amazon Linux AMI, which is supported in all regions and on all current-generation EC2 instance types.  The Amazon Linux AMI supports both PV and HVM mode, as well as both EBS-backed and Instance Store-backed AMIs.

    You can launch this new version of the AMI in the usual ways. You can also upgrade an existing EC2 instance by running the following commands:

    $ sudo yum clean all
    $ sudo yum update

    And then rebooting the instance.

    New Kernel
    A major new feature in this release is the 4.1.7 kernel, which is the most recent long-term stable release kernel. Of particular interest to many customers is the support for OverlayFS in the 4.x kernel series.

    New Features
    The roadmap for the Amazon Linux AMI is driven in large part by customer requests. During this release cycle, we have added a number of features as a result of these requests; hereโ€™s a sampling:

    • Based on numerous customer requests and in order to support joining Amazon Linux AMI instances to an AWS Directory Service directory, we have added Samba 4.1 to the Amazon Linux AMI repositories, available via sudo yum install samba.
    • Numerous customers have asked for PostgreSQL 9.4 and it is now available in our Amazon Linux AMI repositories as a separate package from PostgreSQL 9.2 and 9.3. PostgreSQL 9.4 is available via sudo yum install postgresql94 and the 2015.09 Amazon Linux AMI repositories include PostgreSQL 9.4.4.
    • A frequent customer request has been MySQL 5.6, and we are pleased to offer it in the 2015.09 repositories as a separate package from MySQL 5.1 and 5.5. MySQL 5.6 is available via sudo yum install mysql56 and the 2015.09 Amazon Linux AMI repositories include MySQL 5.6.26.
    • We introduced support for Docker and Go in our 2014.03 AMI, and we continue to follow upstream developments in each. The lead-up to the 2015.09 release included an update to Go 1.4 and to Docker 1.7.1.
    • We already provide Python 2.6, 2.7 (default), and 3.4 in the Amazon Linux AMI, but several customers have also asked for the PyPy implementation of Python. Weโ€™re pleased to include PyPy 2.4 in our preview repository. PyPy 2.4 is compatible with Python 2.7.8 and is installable via sudo yum --enablerepo=amzn-preview install pypy.
    • In our 2015.03 release we added an initial preview of the Rust programming language. Upstream development has continued on this language, and we have updated from Rust 1.0 to Rust 1.2 for the 2015.09 release. You can install the Rust compiler by running sudo yum --enablerepo=amzn-preview install rust.

    The release notes contain a longer discussion of the new features and updated packages, including an updated version of Emacs prepared specially for Jeff in order to ensure timely publication of this blog post!

    โ€” Max Spevack, Development Manager, Amazon Linux AMI.

    PS โ€“ If you enjoy the Amazon Linux AMI offering and would like to work on future versions, let us know!