The Wayback Machine - https://web.archive.org/web/20150905224335/https://docs.aws.amazon.com/ElasticMapReduce/latest/ReleaseGuide/emr-configure-apps.html
Menu
Amazon Elastic MapReduce
Amazon EMR Release Guide (API Version 2009-03-31)

Configuring Applications

You can override the default configurations for applications you install by supplying a configuration object when specifying applications you want installed at cluster creation time. Configuration objects consist of a classification, properties, and optional nested configurations. A classification refers to an application-specific configuration file. Properties are the settings you want to change in that file. You typically supply configurations in a list, allowing you to edit multiple configuration files in one JSON list.

Example JSON for a list of configurations is provided below:

[
  {
    "Classification": "core-site",
    "Properties": {
      "hadoop.security.groups.cache.secs": "250"
    }
  },
  {
    "Classification": "mapred-site",
    "Properties": {
      "mapred.tasktracker.map.tasks.maximum": "2",
      "mapreduce.map.sort.spill.percent": "90",
      "mapreduce.tasktracker.reduce.tasks.maximum": "5"
    }
  }
]

The classification usually specifies the file name that you want modified. An exception to this is the deprecated bootstrap action configure-daemons, which is used to set environment parameters such as --namenode-heap-size. Now, options like this are subsumed into the hadoop-env and yarn-env classifications with their own nested export classifications. Another exception is s3get, which was used to place a customer EncryptionMaterialsProvider object on each node in a cluster for use in client-side encryption. An option was added to the emrfs-site classification for this purpose.

An example of the hadoop-env classification is provided below:

[
  {
    "Classification": "hadoop-env",
    "Properties": {
      
    },
    "Configurations": [
      {
        "Classification": "export",
        "Properties": {
          "HADOOP_DATANODE_HEAPSIZE": "2048",
          "HADOOP_NAMENODE_OPTS": "-XX:GCTimeRatio=19"
        },
        "Configurations": [
          
        ]
      }
    ]
  }
]

An example of the yarn-env classification is provided below:

[
  {
    "Classification": "yarn-env",
    "Properties": {
      
    },
    "Configurations": [
      {
        "Classification": "export",
        "Properties": {
          "YARN_RESOURCEMANAGER_OPTS": "-Xdebug -Xrunjdwp:transport=dt_socket"
        },
        "Configurations": [
          
        ]
      }
    ]
  }
]

Bootstrap actions that were previously used to configure Hadoop and other applications are replaced by configurations. The following tables give classifications for components and applications and corollary bootstrap actions for each. If the classification matches a file documented in the application project, see that respective project documentation for more details.

Hadoop

FilenameAMI Version Bootstrap ActionRelease Label Classification
core-site.xml configure-hadoop -c core-site
log4j.properties configure-hadoop -l hadoop-log4j
hdfs-site.xml configure-hadoop -s hdfs-site
mapred-site.xml configure-hadoop -m mapred-site
yarn-site.xml configure-hadoop -y yarn-site
httpfs-site.xml configure-hadoop -t httpfs-site
capacity-scheduler.xml configure-hadoop -z capacity-scheduler
hadoop-env.sh configure-daemons --client-optshadoop-env
httpfs-env.sh n/a httpfs-env
mapred-env.sh n/a mapred-env
yarn-env.sh configure-daemons --resourcemanager-optsyarn-env

Spark

FilenameAMI Version Bootstrap ActionRelease Label Classification
spark-defaults.confn/aspark-defaults
spark-env.shn/aspark-env
log4j.propertiesn/aspark-log4j

Hive

FilenameAMI Version Bootstrap ActionRelease Label Classification
hive-env.shn/ahive-env
hive-site.xmlhive-script --install-hive-site ${MY_HIVE_SITE_FILE}hive-site
hive-exec-log4j.propertiesn/ahive-exec-log4j
hive-log4j.propertiesn/ahive-log4j

Pig

FilenameAMI Version Bootstrap ActionRelease Label Classification
pig.propertiesn/apig-properties
log4j.properties n/apig-log4j

EMRFS

FilenameAMI Version Bootstrap ActionRelease Label Classification
emrfs-site.xmlconfigure-hadoop -eemrfs-site
n/as3get -s s3://custom-provider.jar -d /usr/share/aws/emr/auxlib/emrfs-site (with new setting fs.s3.cse.encryptionMaterialsProvider.uri)

The following settings do not belong to a configuration file but are used by Amazon EMR to potentially set multiple settings on your behalf.

Amazon EMR-curated Settings

ApplicationRelease Label ClassificationValid PropertiesWhen To Use
SparksparkmaximizeResourceAllocationConfigure executors to utilize maximum resources of each node

Example Supplying a Configuration in the Console

To supply a configuration, you navigate to the Create cluster page and choose Edit software settings. You can then enter the configuration directly (in JSON or using shorthand syntax demonstrated in shadow text) in the console or provide a Amazon S3 URI for a file with JSON Configurations object.


Example Supplying a Configuration Using the CLI

You can provide a configuration to create-cluster by supplying a path to a JSON file stored locally or in Amazon S3:

aws emr create-cluster --release-label emr-4.0.0 --instance-type m3.xlarge --instance-count 2 --applications Name=Hive --configurations https://s3.amazonaws.com/mybucket/myfolder/myConfig.json

Example Supplying a Configuration Using the Java SDK

The following program excerpt shows how to supply a configuration using the AWS SDK for Java:


	Application hive = new Application();
		hive.withName("Hive");

	Map<String,String> hiveProperties = new HashMap<String,String>();
		hiveProperties.put("hive.join.emit.interval","1000");
		hiveProperties.put("hive.merge.mapfiles","true");
	    
	Configuration myHiveConfig = new Configuration()
		.withClassification("hive-site")
		.withProperties(hiveProperties);

	RunJobFlowRequest request = new RunJobFlowRequest()
		.withName("Create cluster with ReleaseLabel")
		.withReleaseLabel("emr-4.0.0")
		.withApplications(hive)
		.withConfigurations(myHiveConfig)
		.withServiceRole("EMR_DefaultRole")
		.withJobFlowRole("EMR_EC2_DefaultRole")
		.withInstances(new JobFlowInstancesConfig()
			.withEc2KeyName("myKey")
			.withInstanceCount(1)
			.withKeepJobFlowAliveWhenNoSteps(true)
			.withMasterInstanceType("m3.xlarge")
			.withSlaveInstanceType("m3.xlarge")
		);