如何做大表和大表的关联?
作者:互联网
如何做大表和大表的关联? 对于大表和大表的关联: 1.reducejoin可以解决关联问题,但不完美,有数据倾斜的可能,如前所述。 2.思路:将其中一个大表进行切分,成多个小表再进行关联。
package com;
import org.apache.commons.lang.StringUtils;
import org.apache.hadoop.io.LongWritable;
import org.apache.hadoop.io.NullWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Mapper;
import java.io.BufferedReader;
import java.io.FileInputStream;
import java.io.IOException;
import java.io.InputStreamReader;
import java.util.HashMap;
import java.util.Map;
public class MapJoinMapper extends Mapper<LongWritable, Text, Text, NullWritable> {
Map<String, String> dictMap = new HashMap<>();
Text k = new Text();
protected void setup(Context context) throws IOException, InterruptedException {
String path = context.getCacheFiles()[0].getPath();
BufferedReader br = new BufferedReader(new InputStreamReader(new FileInputStream(path)));
String line;
while (StringUtils.isNotEmpty(line = br.readLine())) {
String[] fields = line.split(",");
dictMap.put(fields[0], fields[1]+","+fields[2]);
}
br.close();
}
protected void map(LongWritable key, Text value, Context context) throws IOException, InterruptedException {
String orderLine = value.toString();
String[] fields = orderLine.split(",");
更多内容请见原文,文章转载自:https://blog.csdn.net/qq_44594249/article/details/96970999
标签:java,String,fields,关联,做大表,io,new,import,大表 来源: https://www.cnblogs.com/xiaolongxia1922/p/15516149.html