Spark provides programmers with an application programming interface centered on a data structure called the resilient distributed dataset (RDD), a read-only multiset of data items distributed over a cluster of machines, that is maintained in a fault-tolerant way. It was developed in response to limitations in the MapReduce cluster computing paradigm, which forces a particular linear dataflowstructure on distributed programs: MapReduce programs read input data from disk, map a function across the data, reduce the results of the map, and store reduction results on disk. Spark's RDDs function as a working set for distributed programs that offers a (deliberately) restricted form of distributed shared memory
The BlueData EPIC™ (Elastic Private Instant Clusters) software platform solves the infrastructure challenges and limitations that can slow down and stall Big Data deployments. With EPIC software, you can spin up Hadoop and Spark clusters – with the data and analytical tools that your data scientists need – in minutes rather than months.
Leveraging the power of containers, BlueData EPIC makes it easier, faster, and more cost-effective to deploy Big Data infrastructure and applications—including Hadoop, Spark, Kafka, Cassandra, and more— whether on-premises or in the public cloud. Your data scientists and analysts can use the tools they prefer. You can run it with any shared storage environment, so you don’t have to move your data. And it delivers the enterprise-grade security and governance that your IT teams require.
With the BlueData EPIC software platform, you can provide a highly flexible and secure Big-Data-as-a-Service environment to enable faster time-to-insights and faster time-to-value – now available either on-premises or on AWS.
Cloudera Distribution for Hadoop is the world's most complete, tested, and popular distribution of Apache Hadoop and related projects. CDH is 100% Apache-licensed open source and is the only Hadoop solution to offer unified batch processing, interactive SQL, and interactive search, and role-based access controls. More enterprises have downloaded CDH than all other such distributions combined.
NASA JPL, UC Berkeley AMPLab, Amazon, eBay, Yahoo!, UC Santa Cruz, TripAdvisor, Taboola, Agile Lab, Art.com, Baidu, Alibaba Taobao, EURECOM, Hitachi Solutions
The Advisory Board Company, AIG, Attunity, Comcast, conEdison, HGST, Intellisoft, John Hopkins, Tarleton State University, National Supercomputing Center, Orange, Quantres
37signals, Adconion,adgooroo, Aggregate Knowledge, AMD, Apollo Group, Blackberry, Box, BT, CSC