阅读背景:

一个基于大事件的表或多个表?蜂巢表设计考虑

来源:互联网 

I am working on a problem where we have lot of different events coming from different sources and these events have 60% fields common. So, with that said, I initially started with created individual tables for each event and now see that there can many events and almost 60% data fields are same among these events, I am thinking of create one event table that will have columns for all events and I am going to add a type column in this table which will let my spark jobs pick events relevant to them. This table is a Hive external table, and spark jobs will load data into it by processing a staging json table. I am working on a problem where we have lot of




你的当前访问异常,请进行认证后继续阅读剩余内容。

分享到: