AWS Official Blog

  • AWS Storage Update โ€“ New Lower Cost S3 Storage Option & Glacier Price Reduction

    by Jeff Barr | on | in Amazon Glacier, Amazon S3, Price Reduction | | Comments

    Like all AWS services, the Amazon S3 team is always listening to customers in order to better understand their needs. After studying a lot of feedback and doing some analysis on access patterns over time, the team saw an opportunity to provide a new storage option that would be well-suited to data that is accessed infrequently.

    The team found that many AWS customers store backups or log files that are almost never read. Others upload shared documents or raw data for immediate analysis. These files generally see frequent activity right after upload, with a significant drop-off as they age. In most cases, this data is still very important, so durability is a requirement. Although this storage model is characterized by infrequent access, customers still need quick access to their files, so retrieval performance remains as critical as ever.

    New Infrequent Access Storage Option
    In order to meet the needs of this group of customers, we are adding a new storage class for data that is accessed infrequently. The new S3 Standard โ€“ Infrequent Access (Standard โ€“ IA) storage class offers the same high durability, low latency, and high throughput of S3 Standard. You now have the choice of three S3 storage classes (Standard, Standard โ€“ IA, and Glacier) that are designed to offer 99.999999999% (eleven nines) of durability.โ€Ž  Standard โ€“ IA has an availability SLA of 99%.

    This new storage class inherits all of the existing S3 features that you know (and hopefully love) including security and access management, data lifecycle policies, cross-region replication, and event notifications.

    Prices for Standard โ€“ IA start at $0.0125 / gigabyte / month (one and one-quarter US pennies), with a 30 day minimum storage duration for billing, and a $0.01 / gigabyte charge for retrieval (in addition to the usual data transfer and request charges). Further, for billing purposes, objects that are smaller than 128 kilobytes are charged for 128 kilobytes of storage. We believe that this pricing model will make this new storage class very economical for long-term storage, backups, and disaster recovery, while still allowing you to quickly retrieve older data if necessary.

    You can define data lifecycle policies that move data between Amazon S3 storage classes over time. For example, you could store freshly uploaded data using the Standard storage class, move it to Standard โ€“ IA 30 days after it has been uploaded, and then to Amazon Glacier after another 60 days have gone by.

    The new Standard โ€“ IA storage class is simply one of several attributes associated with each S3 object. Because the objects stay in the same S3 bucket and are accessed from the same URLs when they transition to Standard โ€“ IA, you can start using Standard โ€“ IA immediately through lifecycle policies without changing your application code. This means that you can add a policy and reduce your S3 costs immediately, without having to make any changes to your application or affecting its performance.

    You can choose this new storage class (which is available today in all AWS regions) when you upload new objects via the AWS Management Console:

    You can set up lifecycle rules for each of your S3 buckets. Hereโ€™s how you would establish the policies that I described above:

    These functions are also available through the AWS Command Line Interface (CLI), the AWS Tools for Windows PowerShell, the AWS SDKs, and the S3 API.

    Hereโ€™s what some of our early users have to say about S3 Standard โ€“ Infrequent Access:

    โ€œFor more than 13 years, SmugMug has provided unlimited storage for our customerโ€™s priceless photos. With many petabytes of them stored on Amazon S3, itโ€™s vital that customers have immediate, instant access to any of them at a momentโ€™s notice โ€“ even if they havenโ€™t been viewed in years. Amazon S3 Standard โ€“ IA offers the same high durability and performance as Amazon S3 Standard so we can continue to deliver the same amazing experience for our customers even as their cameras continue to shoot bigger, higher-quality photos and videos.โ€

    Don MacAskill, CEO & Chief Geek

    SmugMug

    โ€œWe store a ton of video, and in many cases an object in Amazon S3 is the only copy of a userโ€™s video. This means durability is absolutely critical, and so we are thrilled that Amazon S3 Standard โ€“ IA lets us significantly reduce storage costs on our older video objects without sacrificing durability. We also really appreciate how easy it is to start using Amazon S3 Standard โ€“ IA. With a few clicks we set up lifecycle policies that will transition older objects to Amazon S3 Standard โ€“ IA at regular intervals โ€“we donโ€™t have to worry about migrating them to new buckets, or impacting the user experience in any way.โ€

    Brian Kaiser, CTO

    Hudl

    See the S3 Pricing page for complete pricing information on this new storage class.

    Reduced Price for Glacier Storage
    Effective September 1, 2015, we are reducing the price for data stored in Amazon Glacier from $0.01 / gigabyte / month to $0.007 / gigabyte / month. As usual, this price reduction will take effect automatically and you need not do anything in order to benefit from it. This price is for the US East (Northern Virginia), US West (Oregon), and Europe (Ireland) regions; take a look at the Glacier Pricing page for full information on pricing in other regions.

    โ€” Jeff;

  • Docker Trusted Registry โ€“ Now in the AWS Marketplace

    by Jeff Barr | on | in AWS Marketplace, EC2 Container Service | | Comments

    During my trip to the AWS Loft earlier this month, I spoke to 8 startups for the AWS Podcast.  Almost all of them told me that they are making use of Docker on AWS, either directly or via Amazon EC2 Container Service. They love the flexibility that it gives them, and appreciate the ease with which they can move from development (often on a laptop) to test, and then on to production while remaining highly confident that their code and their configurations will work as expected in each environment.

    In order to enable a very wide variety of use cases, we are making the Docker Trusted Registry (DTR) available in the AWS Marketplace. You can launch it on an EC2 instance in order to create a private registry.

    This new offering supports the popular laptop-to-cloud workflow by giving you a central, highly accessible location to store and manage your Docker images for deployment in your chosen on-premises or cloud environment. You can create custom access control levels and use them to regulate access to the images in your registry. You can require the use of SSL certificates or LDAP entries, and you can take advantage of all of the network access controls that are part of the Virtual Private Cloud (VPC).

    To learn more about the configuration options that are available to you, read the post New AWS Support for Commercially-Supported Docker Applications: Docker Trusted Registry and Docker Engine on the AWS Partner Network Blog.

    โ€” Jeff;

  • Alert Logic Cloud Insight โ€“ Product Tour

    by Jeff Barr | on | in Guest Post, Security | | Comments

    I love to see all of the cool products and services that the Members of the AWS Partner Network (APN) build and bring to market. In the guest post below, my colleague Shawn Anderson takes you on a tour of Alert Logicโ€™s new Cloud Insight product.

    โ€” Jeff;


    In August, Alert Logic introduced Alert Logic Cloud Insight, which identifies vulnerabilities in operating systems and applications running on EC2 instances and configuration issues with AWS accounts and services. This product discovers and evaluates an AWS environment using data provided by EC2, Virtual Private Cloud, Auto-Scaling, Elastic Load Balancing, IAM, and RDS APIs. Currently, Alert Logic is offering a 30-day free trial of Cloud Insight.

    To begin using Cloud Insight you first login to the Cloud Insight web portal and give Cloud Insight access to your AWS environment via an IAM role. There are step-by-step instructions provided in the product describing how you set up this access:

    Cloud Insight will automatically discover all of the hosts and services associated with your AWS environment. Cloud Insight then automatically creates a dedicated security subnet in your VPC and launches a virtual Alert Logic appliance in the subnet. Within a few minutes you will see the results of the discovery process in the topology view:

    The topology view shows the relationship between your  AWS assets, The relationships (lines between assets) are updated dynamically as your AWS environment changes. To complete your setup, you select the assets you want to be part of Cloud Insightโ€™s continuous assessments. You can choose to protect an entire region, VPC, or subnet. You can make adjustments to this scope at any time.

    Once you finish this step, Cloud Insight is up and running. It will continuously scan your assets and audit your environment configuration, and identify vulnerabilities and configuration issues it encounters. In the topology view you can see where the issues were discovered, color-coded for severity:

    By accessing the Remediation page, you can see a list of prioritized remediation actions that will address the identified vulnerabilities and configuration issues. The prioritization of actions is based on contextual analysis using a proprietary methodology:

    By taking these steps you can see that, for example, an upgrade to one Apache HTTP_server image addresses several vulnerabilities discovered in the environment:

    When a remediation action is completed, mark it complete and Cloud Insight will rescan the impacted hosts to verify that the vulnerability has been eliminated.

    Cloud Insight is well suited for a security analyst who wants to identify critical exposures in their environment.  Additionally Cloud Insight is accessible via APIs meaning that you could incorporate it into a continuous deployment program.  For more information on Cloud Insight you can visit Alert Logicโ€™s website where you can access a few short videos, product documentation, and request your free trial.

    โ€” Shawn Anderson, Global Ecosystem Alliance Lead, AWS Partner Network

  • New Spot Fleet Option โ€“ Distribute Your Fleet Across Multiple Capacity Pools

    by Jeff Barr | on | in Amazon EC2 | | Comments

    Last week I spoke to a technically-oriented audience at the Pacific Northwest PHP Conference. As part of my talk, I described cloud computing as a mix of technology and business, and made my point by talking about Spot instances. The audience looked somewhat puzzled at first, but as I explained further I could see their eyes light up as they started to think about the ways that they could save money for their companies by way of creative coding!

    Earlier this year I wrote about the Spot Fleet API, and showed you how to use it to manage thousands of Spot instances with a single call to the RequestSpotFleet function. Today we are introducing a new โ€œallocation strategyโ€ option for that API. This option will allow you to create a Spot fleet that contains instances drawn from multiple capacity pools (a set of instances of a given type, within a particular region and Availability Zone).

    As part of your call to RequestSpotFleet, you can include up to 20 launch specifications. If you make an untargeted request (by not specifying an Availability Zone or a subnet), you can target multiple capacity pools within an AWS region. This gives you access to a lot of EC2 capacity, and allows you to set up fleets that are a good match for your application.

    You can set the allocation strategy to either of the following values:

    • lowestPrice โ€“ This is the default strategy. It will result in a Spot fleet that contains instances drawn from the lowest priced pool(s) specified in your request.
    • diversified โ€“ This is the new strategy, and it must be specified as part of your request. It will result in a Spot fleet that contains instances drawn from all of the pools specified in your request, with the exception of those where the current Spot price is above the On-Demand price.

    This option allows you to choose the strategy that most closely matches your goals for each Spot fleet. The following table can be used as a guide:

    lowestPrice diversified
    Fleet Size Fine for modest-sized fleets. However, a request for a large fleet can affect pricing in the pool with the lowest price. Works well with larger fleets.
    Total Fleet Operating Cost Can be unexpectedly high if pricing in the pool spikes. Should average 70%-80% off of On-Demand over time.
    Consequence of Capacity Fluctuation in a Pool Entire fleet subject to possible interruption and subsequent replenishment. Fraction of fleet (1/Nth of total capacity) subject to possible interruption and subsequent replenishment.
    Application Characteristics Short-running.
    Not time sensitive.
    Long-running.
    Time sensitive.
    Typical Applications Scientific simulations, research computations. Transcoding, customer-facing web servers, HPC, CI/CD.

    If you create a fleet using the diversified strategy and use it to host your web servers, it is a good idea to select multiple pools and to have a fallback option in case all of them become unavailable.

    Diversified allocation works really well in conjunction with the new resource-oriented bidding feature that we launched last month. When you use resource-oriented bidding and specify diversified allocation, each of the capacity pools in your launch specification will include the same number of capacity units.

    To make use of this new strategy, simply include it in your CLI or API-driven request. If you are using the CLI, simply add the following entry to your configuration file:

    "AllocationStrategy": "diversified"

    If you are using the API, specify the same value in your SpotFleetRequestConfigData.

    This option is available now and you can start using it today.

    โ€” Jeff;

  • Elastic Load Balancing Update โ€“ More Ports & Additional Fields in Access Logs

    by Jeff Barr | on | in Amazon EC2, Amazon Elastic Load Balancer | | Comments

    Many AWS applications use Elastic Load Balancing to distribute traffic to a farm of EC2 instances. An architecture of this type is highly scalable since instances can be added, removed, or replaced in a non-disruptive way. Using a load balancer also gives the application the ability to keep on running if an instance encounters an application or system problem of some sort.

    Today we are making Elastic Load Balancing even more useful with the addition of two new features: support for all ports and additional fields in access logs.

    Support for All Ports
    When you create a new load balancer, you need to configure one or more listeners for it. Each listener accepts connection requests on a specific port. Until now, you had the ability to configure listeners for a small set of low-numbered, well-known ports (25, 80, 443, 465, and 587) and to a much larger set of ephemeral ports (1024-65535).

    Effective today, load balancers that run within a Virtual Private Cloud (VPC) can have listeners for any port (1-65535). This will give you the flexibility to create load balancers in front of services that must run on a specific, low-numbered port.

    You can set this up in all of the usual ways: the ELB API, AWS Command Line Interface (CLI) / AWS Tools for Windows PowerShell, a CloudFormation template, or the AWS Management Console. Hereโ€™s how you would define a load balancer for port 143 (the IMAP protocol):

    To learn more, read about Listeners for Your Load Balancer in the Elastic Load Balancing Documentation.

    Additional Fields in Access Logs
    You already have the ability to log the traffic flowing through your load balancers to a location in S3:

    In order to allow you to know more about this traffic, and to give you some information that will be helpful as you contemplate possible configuration changes, the access logs now include some additional information that is specific to a particular protocol. Hereโ€™s the scoop:

    • User Agent โ€“ This value is logged for TCP requests that arrive via the HTTP and HTTPS ports.
    • SSL Cipher and Protocol โ€“ These values are logged for TCP requests that arrive via the HTTPS and SSL ports.

    You can use this information to make informed decisions when you think about adding or removing support for particular web browsers, ciphers, or SSL protocols. Hereโ€™s a sample log entry:

    2015-05-13T23:39:43.945958Z my-loadbalancer 192.168.131.39:2817 10.0.0.1:80 0.000086 0.001048 0.001337 200 200 0 57 "GET https://www.example.com:443/ HTTP/1.1" "curl/7.38.0" DHE-RSA-AES128-SHA TLSv1.2

    You can also use tools from AWS Partners to view and analyze this information. For example, Splunk shows it like this:

    And Sumo Logic shows it like this:

     

    To learn more about access logging, read Monitor Your Load Balancer Using Elastic Load Balancing Access Logs.

    Both of these features are available now and you can start using them today!

    โ€” Jeff;

  • Route 53 Improvements โ€“ Calculated Health Checks and Latency Checks

    by Jeff Barr | on | in Route 53 | | Comments

    Amazon Route 53 is a highly available and scalable Domain Name System (DNS) web service. Route 53 connects user requests to infrastructure running in AWS (EC2 instances, load balancers, and S3 buckets), and can also be used to route users to infrastructure outside of AWS. You can configure Route 53 to perform periodic health checks, and to fail over to alternate endpoints should a check fail. You can also create a monitoring & alerting system for your websites and applications using the health checks.

    Today we are adding two new types of health checks: calculated and latency measurement.

    Calculated Health Checks
    You can now combine the results of multiple Route 53 health checks into a single value using Boolean operations (AND, OR, and NOT). This allows you to create a single health check that combines the results of other checks in a useful way. For example, I have three web sites on the same EC2 instance. Iโ€™ll start by creating health checks for each one:

    Next Iโ€™ll create a calculated health check that represents the overall health of the instance:

    As you can see from the screen shot, I could also choose to create a health check that would report as healthy as long as a certain number (perhaps 2 out of 3) of the other health checks were also healthy. This would avoid false alarms and would allow me to take one site down for maintenance without causing this health check to fail.

    While my example checked three different domains that happened to be running on the same instance, this is not a constraint. You could check multiple website components that happened to be on the same domain but managed by different administrative groups within your organization. Or, you could check multiple domains that supply web services to your app, and then roll up the results into a single check that represents the state of your dependencies.

    Latency Measurement Health Checks
    You can also configure Route 53 to measure and report on metrics that affect latency: TCP connection time, time to first byte, and (for SSL connections) the time to complete the SSL handshake. The first one indicates how long it takes Route 53 to establish a connection to the endpoint; the second one indicates the overall time until the first byte of data is returned. The third measures the time it takes to set up an SSL connection, an operation which involves 2 round-trips.

    The latency measurement is performed as a part of the health check, and (if you check on Latency graphs) reported to CloudWatch.

    Hereโ€™s how you configure latency measurement when you are setting up a health check:

    The results are visible in the Console:

    You can view the results with respect a single AWS region by selecting it from menu (a blended result is shown by default). You can also see the individual metrics (currently 32 per endpoint) in the CloudWatch console:

    Available Now
    The new health checks are available now and you can start using them today. Take a look at the Route 53 Pricing page to learn more about pricing for health checks.

    โ€” Jeff;

  • Moving Past Microsoft Windows Server 2003 End-of-Life Using AWS

    by Jeff Barr | on | in Microsoft Windows | | Comments

    In the guest post below, my colleagues Bryan Nairn and Niko Pamboukas list some options for those of you who are still running your applications on Windows Server 2003.

    โ€” Jeff;


     

    As many of you may already know, on July 14th 2015 Microsoft ended its extended support for Windows Server 2003. Microsoft has published and maintains a support lifecycle for their operating systems to provide clarity on the availability of support for their products.  Once an operating system gets to a certain age, and extended support comes to an end, Microsoft stops issuing security and other updates.

    Twelve years have passed since the original release of Windows 2003 and there are still a large number of businesses running critical applications and workloads on the Windows Server 2003 family of products.  If you are one of these organizations still running Windows 2003 based workloads, you are not alone.  Some industry experts estimate that there are more than 10 million servers running 2003 today.  Some of these workloads are virtualized, however many of them are installed on bare metal.  In many cases these workloads are running on the original hardware and the underlying physical servers are close to the end of their useful life.

    The latest hardware currently available in the market may not necessarily be compatible with Windows Server 2003, thus making your purchasing decisions complex.  Likewise, migration to a newer operating system version will likely require the purchase of new hardware, as the newer system will not necessarily contain all the drivers for the existing hardware.

    This can present challenges for you, and many other organizations like yours, when considering what to do with your Windows Server 2003 infrastructure.  We understand that it takes time to plan and execute a migration, and we are here to help.  Whether you are maintaining 32-bit applications in the cloud, moving to a modern Microsoft Windows Server operating system or rewriting legacy applications, AWS can provide you with production-ready options for migration planning.

    Here are some ideas and resources to help you to assess, plan and execute on your migration strategy for Windows Server 2003.

    Move Your 32-bit Apps to the Cloud
    Itโ€™s a common misconception that you cannot run 32-bit applications in the cloud. Amazon Elastic Compute Cloud (EC2) offers 32-bit instances that you can leverage today. You can start with our 32-bit Windows Server 2003 or Windows Server 2008 Amazon Machine Images (AMIs) or you can use VM Import to bring your own Windows virtual machines images in to EC2. These options give you the breathing room you need to stay on 32-bit while you work on additional migration options for your applications.

    You can also run your 32-bit applications on 64-bit instance types. This will give you additional options, including access to more than 4 GB of memory via PAE. In order to take advantage of this feature, youโ€™ll need to contact AWS Developer Support.

    Migrate to a Modern Operating System
    If you are currently running 32-bit Windows 2003 or 2008 EC2 instances but are ready to migrate, now is the perfect time to get onto the latest version. AWS supports in-place upgrades from Windows 2003 to newer Windows operating systems; you can find the details for how to perform OS upgrades by visiting the EC2 Windows upgrade documentation page. As described in the documentation, this process updates the network driver on the instance so that it can be accessed via Remote Desktop after the OS has been upgraded.

    Modernize Legacy Apps
    When you are ready to start the process of modernizing legacy Windows Server 2003 applications, you can find the right resources for your needs with just a few clicks. AWS has an extensive partner network to help migrate applications to newer versions of Windows. The AWS Windows and .NET Developer Center provides tools, documentation, and code samples. The AWS SDK for .NET (which includes its own library, code samples, and Visual Studio templates) makes it easy to build your applications on Windows and .NET. If you are a Visual Studio user, itโ€™s easy to get started with the SDK using the AWS Toolkit for Visual Studio. You can find more info on the EC2 Developer Resources page. You can also get connected and join the community of developers running Windows and .NET on AWS by visiting our Community Forum or AWS on Github.

    If you need help with the migration of your Windows Server 2003 applications, AWS offers various levels of support including technical documentation, the AWS Support Center, and access to qualified AWS partners. AWS partners specialize in cloud migration services; they are ready to help you assess what applications need to move, identify any risks, gaps and/or modifications needed to migrate smoothly even when migrating unsupported applications without redeploying.

    AWS as a Platform for Your Needs
    In addition to Windows Server 2008 and Windows Server 2012 being more secure and supported software solutions, your move from on-premises infrastructure to the cloud can bring additional benefits: the AWS cloud has been architected to provide a cost effective, flexible and secure cloud computing environment. Since you can provision resources as your business dictates and you only pay for what you use, the cost savings of migrating to AWS can be significant. In addition, Amazon Virtual Private Cloud provides you with an additional layer of security by enabling you to create your own logically isolated networks, which you can provision your resources into. With VPC you can specify your IP range, decide which instances are exposed to the internet and which remain private.

    Start Today
    As I mentioned earlier, nowโ€™s a great time to start the planning process and we are here to help you. We anticipate that you will have questions and may want some help with this, so to get started read our essential Windows Server 2003 FAQ as well as the Windows Servers Server 2003 End-Of-Support page which cover many more details on this transition. We realize that your migration away from Windows Server 2003 can be challenging, hopefully AWS can be there to help ease this transition.

    โ€” Bryan Nairn (Senior Product Manager) and Niko Pamboukas (Senior Product Manager)

  • AWS Week in Review โ€“ September 7, 2015

    by Jeff Barr | on | in Week in Review |

    Letโ€™s take a quick look at what happened in AWS-land last week:

    Monday, September 7
    Tuesday, September 8
    Wednesday, September 9
    Thursday, September 10
    Friday, September 11

    New & Notable Open Source

    • redshift-udfs contains SQL for many helpful Redshift UDFs.
    • cfn-flow is a command-line tool for developing CloudFormation templates and deploying stacks.
    • s3-backup-script tars up local files and folder, dumps a MySQL database, and uploads the resulting files to S3.
    • md5s3stash implements content-addressable storage in S3.
    • PhotoEncryptionInCloud is an Android app that stores encrypted photos in AWS.
    • auto-simple-calculator is an auto-pilot for the AWS Simple Monthly Calculator.
    • aws-cloudwatch-chart is a Node module that draws charts for CloudWatch metrics.
    • awsqr generates QR codes for AWS MFA logins.
    • grails-aws is an AWS plugin for Grails.
    • NFLX-Security-Monkey monitors policy changes and alerts on insecure configurations in an AWS account.

    New Customer Success Stories

    New SlideShare Content

    New YouTube Videos

    New Marketplace Applications

    Upcoming Events

    Upcoming Events at the AWS Loft (San Francisco)

    Upcoming Events at the AWS Loft (New York)

    Help Wanted

    Stay tuned for next week! In the meantime, follow me on Twitter and subscribe to the RSS feed.

    โ€” Jeff;

  • User Defined Functions for Amazon Redshift

    by Jeff Barr | on | in Amazon Redshift | | Comments

    The Amazon Redshift team is on a tear. They are listening to customer feedback and rolling out new features all the time! Below you will find an announcement of another powerful and highly anticipated new feature.

    โ€” Jeff;


    Amazon Redshift makes it easy to launch a petabyte-scale data warehouse. For less than $1,000/Terabyte/year, you can focus on your analytics, while Amazon Redshift manages the infrastructure for you. Amazon Redshiftโ€™s price and performance has allowed customers to unlock diverse analytical use cases to help them understand their business. As you can see from blog posts by Yelp, Amplitude and Cake, our customers are constantly pushing the boundaries of whatโ€™s possible with data warehousing at scale.

    To extend Amazon Redshiftโ€™s capabilities even further and make it easier for our customers to drive new insights, I am happy to announce that Amazon Redshift has added scalar user-defined functions (UDFs). Using PostgreSQL syntax, you can now create scalar functions in Python 2.7 custom-built for your use case, and execute them in parallel across your cluster.

    Hereโ€™s a template that you can use to create your own functions:

    CREATE [ OR REPLACE ] FUNCTION f_function_name 
    ( [ argument_name arg_type, ... ] )
    RETURNS data_type
    { VOLATILE | STABLE | IMMUTABLE }
    AS $$
      python_program
    $$ LANGUAGE plpythonu;
    

    Scalar UDFs return a single result value for each input value, similar to built-in scalar functions such as ROUND and SUBSTRING. Once defined, you can use UDFs in any SQL statement, just as you would use our built-in functions.

    In addition to creating your own functions, you can take advantage of thousands of functions available through Python libraries to perform operations not easily expressed in SQL. You can even add custom libraries directly from S3 and the web. Out of the box, Amazon Redshift UDFs come integrated with the Python Standard Library and a number of other libraries, including:

    • NumPy and SciPy, which provide mathematical tools you can use to create multi-dimensional objects, do matrix operations, build optimization algorithms, and run statistical analyses.
    • Pandas, which offers high level data manipulation tools built on top of NumPy and SciPy, and that enables you to perform data analysis or an end-to-end modeling workflow.
    • Dateutil and Pytz, which make it easy to manipulate dates and time zones (such as figuring out how many months are left before the next Easter that occurs in a leap year).

    UDFs can be used to simplify complex operations. For example, if you wanted to extract the hostname out of a URL, you could use a regular expression such as:

    SELECT REGEXP_REPLACE(url, '(https?)://([^@]*@)?([^:/]*)([/:].*|$)', โ€˜\3') FROM table;
    

    Or, you could import a Python URL parsing library, URLParse, and create a function that extracts hostnames:

    CREATE FUNCTION f_hostname(url VARCHAR)
    RETURNS varchar
    IMMUTABLE AS $$
    import urlparse
    return urlparse.urlparse(url).hostname
    $$ LANGUAGE plpythonu;
    

    Now, in SQL all you have to do is:

    SELECT f_hostname(url) 
    FROM table;
    

    As our customers know, Amazon Redshift obsesses about security. We run UDFs inside a restricted container that is fully isolated. This means UDFs cannot corrupt your cluster or negatively impact its performance. Also, functions that write files or access the network are not supported. Despite being tightly managed, UDFs leverage Amazon Redshiftโ€™s MPP capabilities, including being executed in parallel on each node of your cluster for optimal performance.

    To learn more about creating and using UDFs, please see our documentation and a detailed post on the AWS Big Data blog. Also, check out this how-to guide from APN Partner Looker. If youโ€™d like to share the UDFs youโ€™ve created with other Amazon Redshift customers, please reach out to us at redshift-feedback@amazon.com. APN Partner Periscope has already created a number of useful scalar UDFs and published them here.

    We will be patching your cluster with UDFs over the next two weeks, depending on your region and maintenance window setting. The new cluster version will be 1.0.991. We know youโ€™ve been asking for UDFs for some time and would like to thank you for your patience. We look forward to hearing from you about your experience at redshift-feedback@amazon.com.

    โ€” Tina Adams, Senior Product Manager

  • AWS Podcasts โ€“ Legion Analytics, Bohemian Guitars, Remind, Remeeting

    by Jeff Barr | on | in AWS Podcast | | Comments

    Earlier this month I spent two exceptionally pleasant days at the AWS Loft in San Francisco. While I was there I sat down with a number of startups and recorded their stories. These stories are part of a new Intel Startup Spotlight series that Iโ€™ll be focusing on in the coming weeks and months. As an experiment, I am releasing four podcasts simultaneously.

    In the past, I spent a lot of time editing the podcast to remove some pauses and some background noise. This was very time consuming and made for a marginally better product. In an attempt to get these stories to you on a more timely basis, I am now relaxing my standards and presenting the content to you on a more-or-less as-recorded basis. Because these interviews were recorded in the basement of the Loft, you will hear the occasional footstep or siren, along with the rumblings as the BART trains pass by underneath Market Street.

    On Monday, August 31 I spoke with the following startups:

    • Legion Analyticsโ€“ Automated lead generation.
    • Bohemian Guitars โ€“ Next-generation electric guitars.
    • Remind โ€“ Messaging for teachers, parents, and students.
    • Remeeting โ€“ Meeting recording and analytics.

    One of the big takeaways from two days interviews with these startups (apart from the obvious creativity and intensity) is just how quickly containers have become an integral aspect of the systems that these developers are building. They spoke glowingly of Docker and make good use of Amazon EC2 Container Service to encapsulate applications and to support quick-turn development practices.

    Here are the episodes and the show notes (the โ€œEpisodeโ€ links go directly to the MP3 files):

    Episode 107 โ€“ Legion Analytics
    For Episode 107, I interviewed Jamasen Rodriguez (co-founder and CEO. below in the center) and Sinan Ozdemir (co-founder and CTO, below at left) of Legion Analytics to learn more about how they built an automated lead generation platform on AWS using AWS Elastic Beanstalk, Amazon Simple Storage Service (S3), and other services. Jamasen and Sinan offered a free trail of their product; email them (yourfriends@legionanalytics.com) for more info.

    Episode 108 โ€“ Bohemian Guitars
    For Episode 108, I interviewed Adam Lee (Co-founder) of Bohemian Guitars. He told me how he and his brother were inspired by the musicians of South Africa, who repurposed discarded materials into musical instruments, starting from a ping-pong table in their parentsโ€™ basement. Adam shared some tips that will be useful to anyone who wants to run a successful KickStarter or Indiegogo campaign.

    Episode 109 โ€“ Remind
    For Episode 109 I spoke with engineers Mike Barrett and Eric Holmes of Remind.com to learn more about how they provide teachers with a better way to communicate with students and their parents. They have 25 million users and have sent over 2 billion messages to date, peaking at about 85,000 requests per minute to the API that is consumed by their mobile clients. The site runs on AWS (atop the Empire PaaS that was also built at Remind) and is backed by dozens of microservices. Events are dumped in to a โ€œreally bigโ€ Redshift cluster for analysis.

    Episode 110 โ€“ Remeeting
    For Episode 110 I spoke with Arlo Faria, founder of Remeeting.com . The voice recorder app runs on iOS and Android devices and uploads the raw (but compressed) audio to AWS for post-meeting analysis. As Arlo puts it, โ€œthe magic happens afterward.โ€ Using speech technology that Arlo and his co-founder developed at UC Berkeley, they identify individual speakers and isolate phrases, and present the result in a structured, color-coded form.

    In the Works
    I will be publishing the second day of Intel Startup Spotlight interviews shortly. After AWS re:Invent, I plan to interview even more startups in Seattle and Portland (Oregon). Stay tuned for more info!

    โ€” Jeff;

    PS โ€“ Special thanks are due to my colleague Gloria Kim for settings up the interviews and for taking the pictures.