Showing posts with label Hadoop. Show all posts
Showing posts with label Hadoop. Show all posts

Friday, May 16, 2014

How to drive a project on NoSQL, Big data, Elasticsearch, MongoDB, Hadoop, and other such technologies

I'm reading: How to drive a project on NoSQL, Big data, Elasticsearch, MongoDB, Hadoop, and other such technologiesTweet this !
I am authoring this blog after quite a long break from blogging. Once one gets married, promoted in the organization at the same time, and made responsible for more than 20+ projects as the Lead Architect for a portfolio, it's not easy to catch up with blogging. 

These days, I work on projects spanning technologies like Sharepoint 2013, .NET, jQuery, SQL Server, SSIS, SSAS, SSRS, Powerpivot, Powerview, Mobile web apps using Bootstap and jQueryMobile, Native apps using iOS xCode, and NoSQL based technologies like Elasticsearch and MongoDB. Working as a solution architect with a broad range of projects and technologies is like working as a chef in a kitchen. I get to mix and merge various technology combinations, to create various solution recipes that cater to project requirements. The only exception is bad recipes are not tolerated easily as significant cost is involved based on a architect's decision.

I have spent my career working with technologies that were predominantly from Microsoft space. But the world is changing, and so are the focus on technologies. I have been taking a lot of personal interest in studying more on the NoSQL based technologies that can tap intelligence from unstructured data as well as big data.

One of the biggest traits that many developer or architect generally have is the typical punch line "I can't learn by reading, I need hands-on experience of the technology I need to manage". If you are working with a multi-national organization, it's not that easy to land into a project where neither you would have an experience in the driving technology, and in most cases neither the organization would have any experience too. When organizations don't find or recognize use-cases for any particular technology, if you try to push or propose the technology, it would be seen as you are trying to sell the technology and it's a solution in search of a problem. 

So the big question is, how to bag an entry ticket into the NoSQL world and drive a project using NoSQL technologies ?

Some of the initiatives that can help professionals seeking to build competency in NoSQL as well as intending to drive NoSQL based projects, can consider the following points:

1) Setup a personal lab: Virtualization has made is easy to create a VM. Most of the NoSQL technologies require very modest resources (like 2 GB RAM and single core), to run the software. This can be a starting playground to start practicing the technology.

2) Join the global community: Platforms like Github and Stackoverflow have lot of community projects and real-life questions. By being an active observer as well as participant of these platforms, one can mature on the technology very fast as well as make oneself globally visible as an active professional in the technology of choice.

3) Create a community within your organization: Organizations feel comfortable in adopting technologies, which can be easily managed by the pool of people available in the organization. If you one of the few ones having grip on the technology, you may classify yourself in the niche bracket, but that does not increase organizations confidence to deal in the technology. To deal with this issue, you should conduct various awareness sessions to bring people are various levels up to speed with technology, and create a community of practice in the organization.

4) Pursue a professional training: Post you have been able to successfully pursue points 2 and 3, you can confidently ask for a budget from the organization to pursue professional training on the subject. Everyone's pocket might not allow to pursue training from one's own pocket !!

5) Develop and publish POCs: Confidence to adopt a technology and confidence in a professionals ability to manage a technology, is reflected by the professionals ability to justify the use-case for technology. Identifying use-cases and justifying through POCs are the best means for the same.

By following these 5 steps, I believe that one can establish oneself as well as one's organization in a position to make an entry in the NoSQL world. Let me know what you think.

Sunday, September 23, 2012

Hadoop tools for SSIS, SSRS and SSAS like Integration, Reporting and Analytics

I'm reading: Hadoop tools for SSIS, SSRS and SSAS like Integration, Reporting and AnalyticsTweet this !
Data hosting, processing and reporting is dramatically changing on a variety of extremely different platforms than ever. With the emergence of NoSQL and Big Data, systems such as Hadoop host unimaginable volumes of data. Google is soon to hit 1 Billion Android device activations. US and China collectively contributes to almost 300 million iOS + Android activations. Sourcing data from systems like Hadoop, mashing it up with relational data sources and provisioning reporting and analytics on the most aggressively growing platforms like Android and iOS is not an easy job, leave apart the complexity, cost and skills involved in the process.

Recently I have seen quite a couple of SQL Server and MS BI related blogs writing about how to write code for HBase, Pig and for other similar sources. Industry matures in terms of developer productivity and user friendliness much aggressively than one knows. Talend - an open source provider of tools for managing Big data, provides a tool called Talend Open Studio for Big Data. Its a GUI based data integration tool like SSIS. Behind the scenes this tool generates code for Hadoop Distributed File System (HDFS), Pig, Hbase, Sqoop and Hive. This kind of tools really take Hadoop and Big Data to a extremely wide user-base.


After you have the ways to build a high-way to a mountain of data-source, the immediate need is to make meaning of these data. One of the front-runners of data visualization and analytics, Tableau, provides way to create ad-hoc visualizations from extracts of data from Hadoop clusters or straight live from the Hadoop clusters. Creating visualization from in-memory data and staging extract of data from Hadoop clusters into relational databases and creating visualizations from the same; both are facilitated by Tableau.

Other analytics vendor like Snaplogic and Pentaho also provides tools for operating with Hadoop clusters, which does not require developers to write code. Microsoft has an integrated platform for integration, reporting and analytics (in-memory/olap) and an IDE like SSDS (formerly BIDS).

If tools similar to Talend and Tableau are integrated into SSIS, SSAS, SSRS, DB Engine and SSDT, then Microsoft is one of the best positioned leaders to take Hadoop to a wide audience in their main-stream business. When platforms like Azure Data Market, Data Quality Services, Master Data Management, StreamInsight, Sharepoint etc join hands with tool and technology support integrated with SQL Sever, it would be an unmatched way to extract intelligence out of Hadoop. Connectors for Hadoop has been the first baby step towards this area. Still lot of maturity in this area is awaited.

Till then look out for existing leaders in this area like Cloudera, MapR, Hortonworks, Apache and GreenPlum for Hadoop distributions and implementation. And for Hadoop tools, software vendors like Talend, Tableau, SnapLogic and Pentaho can provide the required toolset. 

Wednesday, July 04, 2012

Data Integration Services on SQL Azure platform

I'm reading: Data Integration Services on SQL Azure platformTweet this !
They say that any knowledge never goes waste, but in IT parlance this saying can be redefined as any knowledge or DATA never goes waste. In my career till date, my experience has been that any successful business would have lot of external data processing as a part of its business functions. Competitive Sales Intelligence for example is one of the category of data that many IT organizations would keep of processing to defines best analytical insights for their sales team. Companies like Facebook and Google accumulate hoards of data and makes a fortune out of advertising business. But even these giants depend on external data providers for their business functions. Example of one such data provider is Factual, that provides data to Facebook. Monetizing on carefully curated and certified business databases has become a very big business.

Cloud takes this platform of sharing and trading data one step ahead. Windows Azure Marketplace provides DataMarket for the same purpose where applications can share and trade data.

1) Private cloud and public cloud comes into question, as organizations might want to share their datasets but limited to the scope of the organization only. Microsoft codename "Data Hub" claims to provide a flavor of managed self-service enterprise data integration on the cloud, which generally takes a huge team and data centers to serve the same needs of an organization. This platform is expected to provide private data-market to enterprises which can be very interesting in terms of agility and cost-savings.



2) Any sizeable organization would generate and consume lot of internal as well as external data. Integration and sharing of data is implementation of the solution after the source of data has been recognized. Data discovery from within and outside organization for business needs, is a bigger challenge in itself. Microsoft Codename "Data Explorer" can be seen as self-service SSIS on Azure platform. It provides data discovery from the windows azure marketplace as well as provides features for self-service data mashups from a variety of standard data sources. Hadoop is not yet included in the supported data source list, but if its gets included in the future, this platform can reap immense value and can acts organizations private Google blended with SSIS to create self-service data mashups and again publish the same as a source of data using Data Hub.

Power of Hadoop combined with cloud based tools like Data Hub and Data Explorer can generate business for lot of data providers as well as bring immense value to organizations. Also it would enable better use of data and provide cost-savings in enterprise data integration.

 

Wednesday, June 27, 2012

How to use MS BI with Hadoop and Why to use Hadoop with SSAS

I'm reading: How to use MS BI with Hadoop and Why to use Hadoop with SSASTweet this !
IT Professionals who use DBMS, SQL, ETL, Reporting and/or Advanced Reporting, and Analytics consider this as the end of data ecosystems. But this is just mainstream IT sphere in the world of data. I do not intend to emphasize of the potential riding on Hadoop, as there are tons of reference material available for the same. If you want to quick check the direction of wind, you can simply fly a kite, you dont need a satellite weather report. Translating it into plain terms, if you want to get a hint of Hadoop's potential, just google on what data and analytics related companies are upto these days. You would find that database giants like TeraData, Microsoft, Informatica and others are ramping up big efforts to provide support for Hadoop. The big businesses that run on Hadoop are Facebook, Yahoo, LinkedIn, Twitter and others. This suffices to conclude that if you are a vetern opportunist in industry, Hadoop is one of the most promising targets.

The challenge of Hadoop starts with bringing it to mainstream IT, which is mostly warehousing data, reporting it and providing analytics. This methodology generally requires activites like data profiling, data cleansing, ETL, and creating warehouses / marts.

1) From the mainstream database world, Hadoop is a source as well as destination. Its more like a content management system functioning in the the form of a database. Hadoop is a MPP system that can run on parallel nodes reaping peta-byte scale data processing speeds. Cloud is one of the most appropriate infrastructure for the same. Windows Azure was already supporting Hadoop VM installations and now a new offering in underway which is known as Hadoop based services for Windows Azure.

2) Why to use Hadoop when we already have SSAS with the power of BISM ? Well any analytics professional would have this question. Theory does not wet the apetite of a practitioner, so the best answer is a case study video featuring how Klout leverages Hadoop and Microsoft BI Technologies to manage BIG Data.

3) Microsoft has announced a connector SQL Server Connector for Apache Hadoop, its a old news. But if you pay attention to detail it says its Sqoop based which is an open source tool provided by Cloudera which imports data from SQL Database into Hadoop Clusters. It can be a very nice to learn tool to start building your skill stakes in the Hadoop world. Any application would have to pump-in and pump-out data from Hadoop, so import export of data from SQL based databases to Hadoop is an inevitable process.

I plan to make MS BI, Analytics and Visual Business Intelligence coupled with BIG Data and cloud as my new regime. I would be sharing my experiences, thoughts and views on the way through my journey. I have introduced a new section on my blog titled Hadoop, BIG Data and Cloud and added a few useful links under the same. I would be adding more to this section, to keep a watch on the same.

In my views, a rolling stone gathers no moss. I intend to earn the same amount of money and respect and recognition I earn in a year, in a months time. I believe that if one has got a dream like this, one needs to be insightful and embrace the change, be a part of the change and make efforts to change the world thats not ready to change.

For my regular blog readers: My blog has remained silent for around 6 months, and my authoring presence has been going south. To my surprise from the blog statistics I was able to make out that the site visits have remained constant and at times have gone even high that it was ever, even without any activity on my blog. This gives me motivation to keep moving on and being an MVP I also feel an obligation on my shoulders to keep sharing my experiences. The reason for my low authoring activity have been my personal life, and I am turning on the lights of my blog after 6 months straight.

In the time that I took almost a break from blogging, two platforms that I have evidenced influencing the IT Industry limited to the scope of my perspective as a Solutions Architect for Business Intelligence and Analytics, are Android and Hadoop. I would discuss about Android at some other time, this post is about Hadoop.

Monday, October 17, 2011

Excel Services data source in Performancepoint Services 2010, Hadoop Data Source : New breed of data sources

I'm reading: Excel Services data source in Performancepoint Services 2010, Hadoop Data Source : New breed of data sourcesTweet this !
Data is the only currency that is generated every second in IT business and its the only currency that every business was to gather and utilize in the best possible way. BIG data and Unstructured data are creating tsunami sized data related challenges for storage, processing, as well as analysis. In the world of structured as well as unstructured data, more and more newer breeds of data sources are evolving and its good to keep a tab on these evolving breed of data sources.

We earlier heard the announcement related to connectors for Hadoop environments. In the new announcement made recently, Microsoft is now propagating Hadoop in its on-premise and cloud based platform with full integration with its regular line of products ranging from Excel to Business Intelligence stack. Hadoop, Hive, Pig etc are the new terms you would hear now in microsoft parlance too, and with this comes the new breed of data sources. You can read more about this announcement from here.

Even in the world of structured data, if you have a tab on the advancements happening in the Microsoft BI world, you would find newer category of data sources. One such example is Excel Services Data Source in Performancepoint Services 2010. Here's a tutorial on the same on Performancepoint Services Team blog. One service application acts as a data source for another service application, it is a very interesting concept in itself and opens up a new range of possibilities.

With SQL Server Denali, even SSRS would be deployed as a service application when you install it in sharepoint integrated mode. So by the integration theory we just discussed between two service applications, there is also a possibility in the future that PPS scorecards can be used as a data source for SSRS Reports, which has always been the other way round till date. With more variety of data sources, the newer challenge on the horizon is selecting the best way to source data as virtually anything can become source of data !

Tuesday, August 09, 2011

MS BI and Hadoop Integration using Hadoop Connectors for SQL Server and Parallel Data Warehouse to analyze structured and unstructured data

I'm reading: MS BI and Hadoop Integration using Hadoop Connectors for SQL Server and Parallel Data Warehouse to analyze structured and unstructured dataTweet this !
Not-only SQL (No SQL) is ruling the world of unstructured data for data storage, warehousing and analytics, with Hadoop being the most successful and widely used technology. There are two choices you can make when something is gaining immense acceptance: either you can abandon and keep competing with your own league or you can partner with it and extend your reach deeper. Microsoft is without doubt one of the leaders in database management, data warehousing and analytics apart from IBM, Oracle and Teradata, but on structured data only. Microsoft Research is trying to churn out its own set of products to deal with BIG data and unstructured data challenges, using federated databases capable of MPP. But Hadoop has already earned a proven reputation and acceptance in this world of unstructured data.

The good news is that Microsoft is embracing Hadoop environments slowly and adopting a symbiotic policy. No organizations would have exclusively structured data or exclusive unstructured data, it's always a combination of both. Azure platform is already support Hadoop implementations. Recently Microsoft announced an upcoming CTP release of two new Hadoop connectors for SQL Server as well as Parallel Data Warehouse. Many visionary DW players are already offering a hybrid BI implementation that allows to use MapReduce (used to query data from Hadoop environments) and SQL together. With the release of Hadoop connector for SQL Server, its highly probable that SQL Server becomes a source for Hadoop environments rather than vice-versa as the ocean full of unstructured data sits in Hadoop environments which is nowhere in the reach of SQL Server to accomodate.

Still the interoperability facilitated by this connector, would empower SQL Server to extract data of interest from this ocean of data hosted in Hadoop environments, making MS BI stack even more powerful. Database Engines, ETLs as well as OLAP Engines would see bigger challenges than ever when clients start using Hadoop as a source for SQL Server, but my viewpoint is that it would mostly work other way round. These connectors are opening a door to the possibility where SQL Server based databases as well as data warehouses can/would be used in combination with Hadoop and MapReduce, effectively creating new opportunities for the entire ecosystem of database community from clients to technicians.

Its too early to know the taste of the food before you actually taste it, but you can predict about the taste from the flavor, and that's what I am trying to do as of now. You can read the announcement about these connectors from here.

Monday, July 25, 2011

Data Warehousing and Analytics on Unstructured Data / Big Data using Hadoop on Microsoft platform

I'm reading: Data Warehousing and Analytics on Unstructured Data / Big Data using Hadoop on Microsoft platformTweet this !
A major community of data warehousing professionals grow up from the old school of Kimball and Inmon methods of data warehousing. Lots of professionals do boast on virtualization, complex MDX querying, performance tuning OLAP engines and managing data warehouse environments of the size of a few hundred GBs or several TBs, as the most niche and challenging jobs they have on their resume. But there is another world of data warehousing and analytics which most would not have explored, and slowly this revolutionary and emerging wave is reaching SMBs which would effectively challenge the world of data warehousing as we practice today. You might come across a question while reading this post, that what has Microsoft to do with it and answer to this question is towards the end of this post.

Data warehouses and data marts developed using Kimball, Inmon or any hybrid methodology can deal with structured data and have scalability challenges too. Appliance solutions such as Parallel Data Warehouse are Microsoft's candidate to deal with such challenges. Some might think that this is the answer to warehouse largest volume of data and build analytical capabilities on the top of it. But this data volume is just a very small piece of the ecosystem. According to Gartner, enterprise data would grow by 650% in 2014 and 85% of the same would be unstructured data, which is also termed as BIG Data.

Have you ever thought of how organizations like Yahoo, Google, Facebook etc organize their data? Which databases do they use? Whether they have data warehousing and analytics? These organizations have some of the largest data volumes in the world. For example, Facebook is heard to have 12 TB of compressed data added per day and 800 TB of compressed data scanned per day. Can you imagine structuring such volume of data using ETL, storing it in data warehouses, aggregating it using OLAP engines in data marts and extracting analytics out of the same ? To handle such volumes of data for data warehousing and analytics, innovative technologies and infrastructure design are required that can support massively parallel processing, and the one I am talking about is named "Hadoop" which is an open-source distributed computing technology and "Hadoop Distributed File System" which is the storage mechanism for handling unstructured data.

Cloud environments like Amazon are already supporting Hadoop, organizations like Cloudera and IBM are supporting commercial distributions of Hadoop, and a lot of big and famous international business majors are already using Hadoop implementation. The biggest implementation is used by Yahoo with 100,000+ CPUs running on 40,000+ computers running Hadoop. An exhaustive list of organizations using Hadoop can be read from here. Organizations are using Hadoop to implement data warehousing and analytics for purposes like Event Analytics, Click Stream Analytics, Text Analytics and more.

For those who are completely afresh to this part of the world, can go through some very interesting reference material mentioned below:

1) The Google File System

2) Data warehousing and Analytics Infrastructure at Facebook

3) Apache Hadoop Wiki

4) Apache Hadoop MapReduce Implementation at Yahoo

5) Setting up Hadoop on VM

Microsoft is aware of the challenges using unstructured data and Hadoop, and is gearing up slowly for the same.

1) Microsoft Research is developing Project Daytona on Azure platform and Project Dryad, which is perceived by the industry as Microsoft's candidate as an alternative for Apache Hadoop.

2) Those who believe that MDX is the top query language that can deal with huge amount of data from OLAP engines, should check out LINQ to HPC to update their GK.

3) With the increasing popularity and success of Hadoop, Microsoft is also supporting Hadoop on Azure platform. You can get an idea of how to deploy Hadoop cluster on Azure platform from here.

The way Microsoft professionals felt that cloud is something new when Azure was introduced, same would be the case when Microsoft would start supporting Hadoop commercially or introduce a commercial alternative for the same. But neither cloud is a recent invention nor technologies like Hadoop to handle, ware house and analyze unstructred data. In my viewpoint, architects and organizations should develop their readiness to deal with the emerging winds of change and upcoming potential business opportunites that unstructured data can offer.
Related Posts with Thumbnails