We asked business professionals to review the solutions they use. Here are some excerpts of what they said:
"The most valuable feature of this solution is its capacity for processing large amounts of data."
"The solution is very stable."
"I feel the streaming is its best feature."
"The features we find most valuable are the machine learning, data learning, and Spark Analytics."
"The main feature that we find valuable is that it is very fast."
"The processing time is very much improved over the data warehouse solution that we were using."
"The memory processing engine is the solution's most valuable aspect. It processes everything extremely fast, and it's in the cluster itself. It acts as a memory engine and is very effective in processing data correctly."
"AI libraries are the most valuable. They provide extensibility and usability. Spark has a lot of connectors, which is a very important and useful feature for AI. You need to connect a lot of points for AI, and you have to get data from those systems. Connectors are very wide in Spark. With a Spark cluster, you can get fast results, especially for AI."
"Provides deep integration with other Azure resources."
"The most valuable features are the IoT hub and the Blob storage."
"Real-time analytics is the most valuable feature of this solution. I can send the collected data to Power BI in real time."
"I like the IoT part. We have mostly used Azure Stream Analytics services for it"
"The solution has a lot of functionality that can be pushed out to companies."
"When you first start using this solution, it is common to run into memory errors when you are dealing with large amounts of data."
"The solution needs to optimize shuffling between workers."
"When you want to extract data from your HDFS and other sources then it is kind of tricky because you have to connect with those sources."
"We've had problems using a Python process to try to access something in a large volume of data. It crashes if somebody gives me the wrong code because it cannot handle a large volume of data."
"We use big data manager but we cannot use it as conditional data so whenever we're trying to fetch the data, it takes a bit of time."
"I would like to see integration with data science platforms to optimize the processing capability for these tasks."
"The graphical user interface (UI) could be a bit more clear. It's very hard to figure out the execution logs and understand how long it takes to send everything. If an execution is lost, it's not so easy to understand why or where it went. I have to manually drill down on the data processes which takes a lot of time. Maybe there could be like a metrics monitor, or maybe the whole log analysis could be improved to make it easier to understand and navigate."
"Stream processing needs to be developed more in Spark. I have used Flink previously. Flink is better than Spark at stream processing."
"If something goes wrong, it's very hard to investigate what caused it and why."
"There may be some issues when connecting with Microsoft Power BI because we are providing the input and output commands, and there's a chance of it being delayed while connecting."
"It is not complex, but it requires some development skills. When the data is sent from Azure Stream Analytics to Power BI, I don't have the access to modify the data. I can't customize or edit the data or do some queries. All queries need to be done in the Azure Stream Analytics."
"The collection and analysis of historical data could be better."
"The solution offers a free trial, however, it is too short."
"Apache Spark is open-source. You have to pay only when you use any bundled product, such as Cloudera."
"The cost of this solution is less than competitors such as Amazon or Google Cloud."
Spark provides programmers with an application programming interface centered on a data structure called the resilient distributed dataset (RDD), a read-only multiset of data items distributed over a cluster of machines, that is maintained in a fault-tolerant way. It was developed in response to limitations in the MapReduce cluster computing paradigm, which forces a particular linear dataflowstructure on distributed programs: MapReduce programs read input data from disk, map a function across the data, reduce the results of the map, and store reduction results on disk. Spark's RDDs function as a working set for distributed programs that offers a (deliberately) restricted form of distributed shared memory
Apache Spark is ranked 1st in Hadoop with 11 reviews while Azure Stream Analytics is ranked 5th in Streaming Analytics with 5 reviews. Apache Spark is rated 8.6, while Azure Stream Analytics is rated 8.2. The top reviewer of Apache Spark writes "Good Streaming features enable to enter data and analysis within Spark Stream". On the other hand, the top reviewer of Azure Stream Analytics writes "A serverless scalable event processing engine with a valuable IoT feature". Apache Spark is most compared with Spring Boot, AWS Batch, AWS Lambda, SAP HANA and Cloudera Distribution for Hadoop, whereas Azure Stream Analytics is most compared with Databricks, Apache Flink, Apache Spark Streaming, Apache NiFi and Google Cloud Dataflow.
We monitor all Hadoop reviews to prevent fraudulent reviews and keep review quality high. We do not post reviews by company employees or direct competitors. We validate each review for authenticity via cross-reference with LinkedIn, and personal follow-up with the reviewer when necessary.