Skip to main content
CS585/DS503. Big Data Management
Team Presentation (1)
Slides source:
Prof. Mohamed Eltabakh & IBM Almaden
Research Center.
1
Yousef Fadila
Yousef@Fadila.net
Abdulaziz Alajaji
asalajaji@wpi.edu
Worcester Polytechnic Institute
CoHadoop: Flexible Data
Placement and Its Exploitation in
Hadoop
Mohamed Eltabakh
Worcester Polytechnic Institute
• Joint work with: Yuanyuan Tian, Fatma Ozcan, Rainer
Gemulla, Aljoscha Krettek, and John McPherson
• IBM Almaden Research Center
Worcester Polytechnic Institute
What is CoHadoop
• CoHadoop is an extension of Hadoop infrastructure,
where:
─ HDFS accepts hints from the application layer to specify related
files
─ Based on these hints, HDFS tries to store these files on the
same set of data nodes
Example
 Files A and B are related
 Files C and D are related
3
File A File B
File DFile C
CoHadoop
 Files A & B are colocated
 Files C & D are colocated
Hadoop
 Files are distributed blindly
over the nodesCoHadoop System
Worcester Polytechnic Institute
Motivation
• Colocating related files improves the performance of
several distributed operations
─ Fast access of the data and avoids network congestion
• Examples of these operations are:
─ Join of two large files.
─ Use of indexes on large data files
─ Processing of log-data, especially aggregations
4 CoHadoop System
Worcester Polytechnic Institute
Background on HDFS
5 CoHadoop System
 Single namenode and many
datanodes
 Namenode maintains the file
system metadata
 Files are split into fixed sized blocks
and stored on data nodes
 Data blocks are replicated for fault
tolerance and fast access (Default is
3)
Default data placement policy
• First copy is written to the node creating the file (write affinity)
• Second copy is written to a data node within the same rack
• Third copy is written to a data node in a different rack
• Objective: load balancing & fault tolerance
Worcester Polytechnic Institute
Data Colocation in CoHadoop
• Introduce the concept of a locator as an additional
file attribute
• Files with the same locator will be colocated on the
same set of data nodes
Example
 Files A and B are related
 Files C and D are related
6
1 File A 1 File B
5 File D5 File C
Storing Files A, B, C, and D in CoHadoop
11
1 1
5 5
55
CoHadoop System
Worcester Polytechnic Institute
Data Placement Policy in CoHadoop
• Change the block placement policy in HDFS to colocate
the blocks of files with the same locator
• Best-effort approach, not enforced
• Locator table stores the mapping of locators and files
• Main-memory structure
• Built when the namenode starts
• While creating a new file:
• Get the list of files with the same locator
• Get the list of data nodes that store those files
• Choose the set of data nodes which stores the highest number of
files
7 CoHadoop System
Worcester Polytechnic Institute
Example of Data Colocation
8 Locator Table
Block 1
Block 2
File A (1)
Block 1
Block 2
Block 3
File C (1)
An HDFS cluster of
5 Nodes, with 3-way
replication
1
5 file B
file A, file C
A1 A2 A1 A2
A1 A2
C1 C2 C3 C1 C2 C3
C1 C2 C3
B1 B2
B1 B2
B1 B2
D1 D2
D1 D2
D1 D2
Block 1
Block 2
File B (5)
Block 1
Block 2
File D
CoHadoop System
 These files are usually post-processed
files, e.g., each file is a partition
Worcester Polytechnic Institute
Target Scenario: Log Processing
• Data arrives incrementally and continuously in separate files
• Analytics queries require accessing many files
• Study two operations:
─ Join: Joining N transaction files with a reference file
─ Sessionazition: Grouping N transaction files by user id, sort by
timestamp, and divide into sessions
• In Hadoop, these operations require a map-reduce job to
perform
9 CoHadoop System
Joining Un-Partitioned Data (Map-Reduce Job)
10
Dataset A Dataset B Different join keys
HDFS stores data blocks
(Replicas are not shown)
Mapper
M
Mapper
2
Mapper
1
Mapper
3
- Each mapper processes one
block (split)
- Each mapper produces the
join key and the record pairs
Reducer 1 Reducer 2 Reducer N
Reducers perform the
actual join
Shuffling and Sorting Phase
Shuffling and sorting over
the network
Joining Partitioned Data (Map-Only Job)
11
Dataset A Dataset B Different join keys
- Partitions (files) are divided
into HDFS blocks
(Replicas are not shown)
- Blocks of the same partition
are scattered over the nodes
Mapper
2
Mapper
1
Mapper
3
- Each mapper processes an
entire partition from both A & B
- Special input format to read the
corresponding partitions
- Most blocks are read remotely
over the network
- Each mapper performs the join
local
local
local
remote
remote
remote
remote remote
remote
remote
CoHadoop: Joining Partitioned/Colocated Data
(Map-Only Job)
12
Dataset A Dataset B Different join keys
- Partitions (files) are divided into
HDFS blocks
(Replicas are not shown)
- Blocks of the related partitions
are colocated
Mapper
2
Mapper
1
Mapper
3
- Each mapper processes an
entire partition from both A & B
- Special input format to read the
corresponding partitions
- Most blocks are read locally
(Avoid network overhead)
- Each mapper performs the join
All blocks
are local
All blocks
are local
All blocks
are local
Worcester Polytechnic Institute
CoHadoop Key Properties
• Simple: Applications only need to assign the locator file
property to the related files
• Flexible: The mechanism can be used by many
applications and scenarios
─ Colocating joined or grouped files
─ Colocating data files and their indexes
─ Colocating a related columns (column family) in columnar store DB
• Dynamic: New files can be colocated with existing files
without any re-loading or re-processing
13
Worcester Polytechnic Institute
Outline
 What is CoHadoop & Motivation
 Data Colocation in CoHadoop
 Target Scenario: Log Processing
• Related Work
• Experimental Analysis
• Summary
CoHadoop System14
Worcester Polytechnic Institute
Related Work
• Hadoop++ (Jens Dittrich et al., PVLDB, Vol. 3, No. 1, 2010)
─ Creates Trojan join and Trojan index to enhance the performance
─ Cogroups two input files into a special “Trojan” file
─ Changes data layout by augmenting these Trojan files
─ No Hadoop code changes, but static solution, not flexible
• HadoopDB (Azza Abouzeid et al., VLDB 2009)
─ Heavyweight changes to Hadoop framework: data stored in local DBMS
─ Enjoys the benefits of DBMS, e.g., query optimization, use of indexes
─ Disrupts the dynamic scheduling and fault tolerance of Hadoop
 Data no longer in the control of HDFS but is in the DB
• HDFS 0.21: provides a new API to plug-in different data placement
policies
15 CoHadoop System
Worcester Polytechnic Institute
Experimental Setup
• Data Set: Financial transactions data generator, augmented
with accounts table as reference data
─ Accounts records are 50 bytes, 10GB fixed size
─ Transactions records are 500 bytes
• Cluster Setup: 41-node IBM SystemX iDataPlex
─ Each server with two quad-cores, 32GB RAM, 4 SATA disks
─ IBM Java 1.6, Hadoop 0.20.2
─ 1GB Ethernet
• Hadoop configuration:
─ Each worker node runs up to 6 mappers and 2 reducers
─ Following parameters are overwritten
 Sort buffer size: 512MB
 JVM’s reused
 6GB JVM heap space per task
16 CoHadoop System
Worcester Polytechnic Institute
Query Types
• Two queries:
─ Join 7 transactions files with a reference accounts file
─ Sessionize 7 transactions file
• Three Hadoop data layouts:
─ RawHadoop: Data is not partitioned
─ ParHadoop: Data is partitioned, but not colocated
─ CoHadoop: Data is both partitioned and colocated
17 CoHadoop System
Worcester Polytechnic Institute
Data Preprocessing and Loading Time
18 CoHadoop System
 CoHadoop and ParHadoop are almost the same and around 40% of Hadoop++
 CoHadoop incrementally loads an additional file
 Hadoop++ has to re-partition and load the entire dataset when new files arrive
Worcester Polytechnic Institute
Hadoop++ Comparison: Query
Response Time
Join Query: CoHadoop vs. Hadoop++
0
500
1000
1500
2000
2500
3000
70GB 140GB 280GB 560GB 1120GB
Dataset Size
Time(Sec)
Hadoop++
CoHadoop
19 CoHadoop System
 Hadoop++ has additional overhead processing the metadata
associated with each block
Worcester Polytechnic Institute
Join Query: Response Time
Join Query
0
1000
2000
3000
4000
5000
6000
70GB 140GB 280GB 560GB 1120GB
Dataset Size
Time(Sec)
CoHadoop-64M CoHadoop-256M
CoHadoop-512M ParHadoop-64M
ParHadoop-256M ParHadoop-512M
RawHadoop-64M RawHadoop-256M
RawHadoop-512M
20 CoHadoop System
 Savings from ParHadoop and CoHadoop are around 40% and 60%, respectively
Worcester Polytechnic Institute
Fault Tolerance
Recovery from Node Failure
0
5
10
15
20
25
30
35
40
45
64MB 512MB
Block size
Slowdown%
CoHadoop ParHadoop
RawHadoop
21 CoHadoop System
 CoHadoop retains the fault tolerance properties of Hadoop
 Failures in map-reduce jobs are more expensive than in map-only jobs
 Failures under larger block sizes are more expensive than under smaller block sizes
After 50% of the job time, a datanode is killed
Worcester Polytechnic Institute
Summary
 CoHadoop is an extension to Hadoop system to
enable colocating related files
 CoHadoop is flexible, dynamic, light-weight, and
retains the fault tolerance of Hadoop
 Data colocation is orthogonal to the applications
 Joins, indexes, aggregations, column-store files, etc…
 Co-partitioning related files is not sufficient, colocation
further improves the performance
22 CoHadoop System
Worcester Polytechnic Institute
Next Paper
CoHadoop System23
Worcester Polytechnic Institute
Eagle-Eyed Elephant (E3):
Split-Oriented Indexing in
Hadoop
Mohamed Eltabakh
Worcester Polytechnic Institute, MA, USA
Joint work with IBM Almaden, CA, USA
F. Özcan, Y. Sismanis, H. Pirahesh, P. Haas, J.
Vondrak
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Talk Outline
Background and Motivation
E3 System Features
 Indexing and Domain Segmentation
 Materialized Views
 Adaptive Caching
Performance and Evaluation
Related Work & Differences
Summary
25
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
E3 Motivation & Objectives
 Typical Scenarios: Analytical query workloads on Hadoop with
selection predicates
─ Multiple (possibly repeated) queries over the same data set
 No Smart Skipping: No indexing (or split elimination) embedded into
Hadoop
─ Queries scan all the data splits (relevant or not)
 Little Users’ Knowledge: Workloads and data may change
─ Users may not know the query workload in advance or the data schema
26
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
E3 Objectives
27
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
 Discovery-based elimination of irrelevant
splits
 No dependency on physical design, No
data movement or DDL
 Adapt to workload and data changes
Worcester Polytechnic Institute
E3: Highlights
28
• JSON-Based Data Model
• Works on all data types/sources that provide a
mapping to JSON (JSON view of the data)
• Split elimination at I/O layer (InputFormat) before creating
map tasks
• Can be integrated into Jaql
• Can be used in hand-coded map-reduce jobs
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Talk Outline
29
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Background and Motivation
E3 System Features
 Indexing and Domain Segmentation
 Materialized Views
 Adaptive Caching
Performance and Evaluation
Related Work & Differences
Summary
Worcester Polytechnic Institute
1) Split-Level Domain Segmentation
• Applied for all numeric and date attributes
─ One-dimensional clustering to produce multiple
ranges (Reduces false-negative hits)
30
a1 a2 a4a3 a5 a6 a8a7 a9 a10
x
Query Q(x): [a1, a10] contains x
[a1,a2], [a3,a4], [a5,a6], [a7,a8], [a9,a10] do not
contain x
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
2) Coarse-Grained Inverted Index
• Split-level as opposed to record-level
• Inverted index implemented using bitmaps
• Run-Length Encoding for effective compression
31
V1 10001010010100000…
V2 00010000100000001…
Vn 10000010000001100…
Fixed-size = # of splits in the input file
FileA
Split 1 Split 2 Split 3 Split NSplit i
{x, …} {x, …}
{x, …}
{x, …}
(x,{1,2, i})
Query Q(x): Only read splits1, 2, and i
Split-Level
Inverted Index
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Inverted Index Limitations
32
v is
infrequent-
scattered
valueFile A
Split 1 Split 2 Split 3 Split NSplit i
{v, …} {v, …}
{v, …}
{v, …}
{v, …}
{v, …}
{v, …}
(V,{1,2,3, …, i, …, N})
Split-Level
Inverted Index
Query Q(v): Must read all splits containing
value v !
 Inverted Index is of no use for infrequent-scattered values
─ Values appearing in many splits, but few times per split
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
3) Materialized Views
33
File A
Split 1 Split 2 Split 3 Split NSplit i
{v, …} {v, …}
{v, …}
{v, …}
{v, …}
{v, …}
{v, …}
AMV
{v, …}
{v, …}
{v, …}
{v, …}
Split 1 Split M
{v, …}
M << N
Query Q(v): read only M splits
(M << N)
• Build a materialized view
AMV for each file A
• Copy the data records
containing v to AMV
• |AMV| << |A| (in splits)
• At query time, E3 re-directs
Q(v) from A to AMV
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Building the Materialized View
34
• MV is relatively very small  |AMV| ≈ (1%-2%) |A|
• Infrequent-scattered values can be too many 
which v’s to select?
• Modeling as optimization problem:
0-1 Knapsack problem
─ Space constraint: AMV can hold M splits (R records)
─ Each value v has a profit and a cost
 Profit(v) = |Splits(v)| – M
 Cost(v) = |Records(v)|
Select subset of values v to:
Maximize Σprofit(v) | Σcost(v) <= R
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Building the Materialized View:
More Challenges
• Submodular 0-1 Knapsack problem because
─ Selecting v and copying its records to AMV changes the cost of all
other values v’ contained in v’s records
• Naïve greedy algorithm
─ Very expensive to do sorting (profit/cost)
• E3 avoids sorting and ignore cost(v) overlapping
─ Estimates an upper bound K values needed to fill in AMV (over
estimate)
─ Maintain the top K in max-heap (profit/cost for each v).
─ One scan over all dataset  Copy records containing top K
values v until AMV is full.
35
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Talk Outline
36
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Background and Motivation
E3 System Features
 Indexing and Domain Segmentation
 Materialized Views
 Adaptive Caching
Performance and Evaluation
Related Work & Differences
Summary
Worcester Polytechnic Institute
Optimizing Conjunctive Predicates
• Conjunctive predicates can be together
very selective
─ But also harder to optimize (each predicate by itself
may not be selective)
37
File A
Split 1 Split 2 Split 3 Split NSplit 4
{v, …} {v, …}
{v, w, …}
{v, …}
{v, …}
{v, …}
{v, …}
{w, …}
{w, …} {w, …}{w, …}
Query Q(v,w)  read split 3 only
 Index cannot help: splits(v) ∩ splits(w) = {1, 2, 3, …, N}
 Materialized Views cannot help: domain is too large to
enumerate
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Handling “nasty” Value-Pairs
• Too expensive to identify all such value pairs (v, w)
• E3’s Solution: Adaptive cache
─ Only “cache” pairs that are:
 Very nasty (high savings in splits if cached)
 Referenced frequently
 Referenced recently
38
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
4) Adaptive Caching for “nasty” Value-Pairs
• Select the value-pairs based on the observed query
workload
• Given (Q = P1 and P2) over values v and w
─ Compute (splits(v) ∩ splits(w)) from the inverted index
─ Monitor which map tasks return output records  splits(v, w)
─ If |splits(v) ∩ splits(w)| >> |splits(v, w)|, then
 Add (v, w, splits(v, w)) to the cache
39
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
E3’s Cache Replacement Policy
• LRU may perform poorly
─ It does not take savings into account
• SFR (Savings-Frequency-Recency)
Replacement Policy
─ Compute a weight for candidate (v,w):
 Savings in splits: the bigger the saving, the higher the
weight
 Frequency: the more frequently queried, the higher the
weight
 Recency: the more recently queried, the higher the weight
40
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
E3 Computation Flow
41
Map-Phase
(split-level)
Reduce-Phase
(dataset-level)
Range statistics
(v, SplitId,
RecordCount, …)
Inverted Index
Map-Phase
(split-level)
Selected subset of
nasty values
Data split
Final output Final output
Materialized view
Final output
Map-reduce
job
Map-only
job
Need two jobs to
pre-process the
data
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
E3 Computation Flow
Building the Materialized View
• 1) Map-reduce job:
─ Reports v and splits ID, and number of records in that split
containing v;
─ Calculate |splits(v)| and |records(v)| In reduce phase to
execute E3 greedy algorithm > output: list of v to be stored.
• 2) Map-only job:
─ Scans the data, copy records that contain selected v’s to
Amv.
42
Worcester Polytechnic Institute
E3 Query Evaluation
(Putting It All Together)
43
Input Format
1) Read file A & set of predicates P
E3 Wrapper
2) Consult E3’s metadata (A, P)
E3
Metadata
3) Return list of relevant
splits
Or AMV
4) Read AMV
4) Read A,
list of splits
OR
5) Input splits to query
evaluation (map-reduce engine)
>> Ranges & inverted index in
light-weight DB
>> Materialized views are in HDFS
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Talk Outline
44
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Background and Motivation
E3 System Features
 Indexing and Domain Segmentation
 Materialized Views
 Adaptive Caching
Performance and Evaluation
Related Work & Differences
Summary
Worcester Polytechnic Institute
Experimental Setup
• Datasets (800GB)
─ Transaction Processing over XML (TPoX) – Orders
 4 levels of nesting, 181 distinct fields
─ Transaction Processing Council (TPCH) – LineItems
 1 level (no nesting),16 distinct fields
• Cluster
─ 41 nodes cluster: 1 master, and 40 data nodes, 8 cores
─ 160 Mappers and 160 Reducers
─ Block size = 64MB, Replication factor = 2
• Performance
1. Wall clock savings at query time
2. Computation cost of (1) Ranges, (2) Indexes, (3)
Materialized view
3. Storage overhead of (1) Ranges, (2) Indexes, (3)
Materialized view
45
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Query Response Time Savings
• Query: read(hdfs(‘input’))  filter (P1 ^ P2)  count();
─ Equality predicates
• Savings depend on selectivity  up to 20x with E3
optimizations
46
TPOX Saving in Query Time (800GB)
0
100
200
300
400
500
600
Full Scan
(83 Waves)
80% 60% 40% 20% 10% 5% 1%
% of Scanned Waves
Time(Sec)
TPCH Saving in Query Time (800GB)
0
100
200
300
400
500
600
Full Scan
(83 Waves)
80% 60% 40% 20% 10% 5% 1%
% of Scanned Waves
Time(Sec)
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Computation Cost (TPoX)
47
• Costs are shared whenever possible
• Requires ~12 selective queries to redeem the cost
! "
#! ! ! "
$! ! ! "
%! ! ! "
&! ! ! "
' ! ! ! "
( ! ! ! "
) ! ! ! "
*+, - . /"0$! "
1. - 2 . , 3/4"
5, 6. 73. 8"
5, 8. 9"
*+, - . /": "
5, 8. 9"
; <"0#= 4" *+, - . /": """""
; <"
*+, - . /": "; <"
: "5, 8. 9"
!"#$%&'$()%
! "# $%*+%, +# - . *$%' */010(1%&! 2+3456678)%
Size: 507MB
Size: 164GB
Size: 7.5GB
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Computation Cost (TPCH)
48
• Requires ~8 selective queries to redeem the
cost
! "
#! ! ! "
$! ! ! "
%! ! ! "
&! ! ! "
' ! ! ! "
( ) *+, -".$! "
/, +0 , *1-2"
3*4, 51, 6"
3*6, 7"
( ) *+, -"8"
3*6, 7"
9 : ".#; 2" ( ) *+, -"8"""""
9 : "
( ) *+, -"8"9 : "
8"3*6, 7"
!"#$%&'$()%
! "# $%*+%, +# - . *$%' */010(1%&! 2, 3456678)%
Size: 61MB
Size: 41GB
Size: 7.6GB
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute
Summary & Lessons Learned
• Eagle-Eyed Elephant (E3) integrates various
indexing and elimination techniques to
effectively eliminate splits (I/O)
• Up to 20x savings can be achieved using E3
optimizations
• Discovery-based, No DDL or data movement
• Partitioning alone is not enough. Also indexing
alone is not enough
• More complex data  More preprocessing cost
 more queries to redeem the cost
49
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research
Worcester Polytechnic Institute50
E3 System EDBT 2013 Mohamed Eltabakh,WPI IBM
Research