During my 6-year Hadoop adventure, I had an opportunity to work with Big Data technologies at several companies ranging from fast-growing startups (e.g. Spotify) to global corporations and academic institutes. What really amazed me was the difference of how the use-cases were defined, how fast valid solutions were built and how money was spent and […]
We are excited to announce that GetInData becomes the coorganizer (together with our partner Evention) of Big Data Tech Warsaw 2017. The conference will be held in Warsaw (Poland), February 9th, 2017.
In this blog post we share motivation, current status and challenges for our new project, called AirHadoop. AirHadoop follows the sharing economy model and it aims to allow companies to use idle Hadoop clusters that belong to somebody else to temporarily gain more computing power and storage. Shared economy A sharing economy is an economic […]
Few months ago I was working on a project with a lot of geospatial data. Data was stored in HDFS, easily accessible through Hive. One of the tasks was to analyze this data and first step was to join two datasets on columns which were geographical coordinates. I wanted some easy and efficient solution. But […]
Camus, a MapReduce job that loads data from Kafka into HDFS, has a number of time-related configuration settings and assumptions. They control how many messages are consumed from Kafka in each Camus run and where the data is stored in HDFS. I summarize them in this blog post.
The LinkedIn Engineering blog is a great resource of technical blog posts related to building and using large-scale data pipelines with Kafka and its “ecosystem” of tools. In this post I provide several pictures and diagrams (including quotes) that summarise how data pipeline has evolved at LinkedIn over the years. The actual content is based […]
We are excited to announce that GetInData became the coorganizer (together with our partner Evention) of Big Data Technology Summit 2016. The conference will be held in Warsaw, February 24-25th.
Go to Big Data Weekly Quiz #10 to start playing this week’s edition. The quiz covers topic from the last issue of Hadoop Weekly and it contains questions about Spark, Succinct Spark, Zeppelin, S3 and EMR. Remember to share your score on Twitter or Facebook! 🙂
One of our client uses Apache Sentry (incubating) to define and enforce authorization rules to data in a Hadoop cluster. In this blog post, I would like to share my experience in using Sentry 1.4.0 with several tools from Hadoop Ecosystem that come in CDH 5.3 and CDH 5.4.
We are extremely happy to inform that GetInData becomes the official sponsor of Warsaw Hadoop User Group!