Java CombineSequenceFileInputFormat类代码示例

OStack程序员社区-中国程序员成长平台 › 门户 › 编程› Java›Java编程经验

原作者: [db:作者] 来自: [db:来源] 收藏邀请

本文整理汇总了Java中org.apache.hadoop.mapred.lib.CombineSequenceFileInputFormat类的典型用法代码示例。如果您正苦于以下问题：Java CombineSequenceFileInputFormat类的具体用法？Java CombineSequenceFileInputFormat怎么用？Java CombineSequenceFileInputFormat使用的例子？那么恭喜您, 这里精选的类代码示例或许可以为您提供帮助。

CombineSequenceFileInputFormat类属于org.apache.hadoop.mapred.lib包，在下文中一共展示了CombineSequenceFileInputFormat类的2个代码示例，这些例子默认根据受欢迎程度排序。您可以为喜欢或者感觉有用的代码点赞，您的评价将有助于我们的系统推荐出更棒的Java代码示例。

示例1: testFormat

import org.apache.hadoop.mapred.lib.CombineSequenceFileInputFormat; //导入依赖的package包/类
@Test(timeout=10000)
public void testFormat() throws Exception {
  JobConf job = new JobConf(conf);

  Reporter reporter = Reporter.NULL;

  Random random = new Random();
  long seed = random.nextLong();
  LOG.info("seed = "+seed);
  random.setSeed(seed);

  localFs.delete(workDir, true);

  FileInputFormat.setInputPaths(job, workDir);

  final int length = 10000;
  final int numFiles = 10;

  // create a file with various lengths
  createFiles(length, numFiles, random);

  // create a combine split for the files
  InputFormat<IntWritable, BytesWritable> format =
    new CombineSequenceFileInputFormat<IntWritable, BytesWritable>();
  IntWritable key = new IntWritable();
  BytesWritable value = new BytesWritable();
  for (int i = 0; i < 3; i++) {
    int numSplits =
      random.nextInt(length/(SequenceFile.SYNC_INTERVAL/20))+1;
    LOG.info("splitting: requesting = " + numSplits);
    InputSplit[] splits = format.getSplits(job, numSplits);
    LOG.info("splitting: got =        " + splits.length);

    // we should have a single split as the length is comfortably smaller than
    // the block size
    assertEquals("We got more than one splits!", 1, splits.length);
    InputSplit split = splits[0];
    assertEquals("It should be CombineFileSplit",
      CombineFileSplit.class, split.getClass());

    // check each split
    BitSet bits = new BitSet(length);
    RecordReader<IntWritable, BytesWritable> reader =
      format.getRecordReader(split, job, reporter);
    try {
      while (reader.next(key, value)) {
        assertFalse("Key in multiple partitions.", bits.get(key.get()));
        bits.set(key.get());
      }
    } finally {
      reader.close();
    }
    assertEquals("Some keys in no partition.", length, bits.cardinality());
  }
}

开发者ID:naver，项目名称:hadoop，代码行数:56，代码来源:TestCombineSequenceFileInputFormat.java

示例2: setUpMultipleInputs

import org.apache.hadoop.mapred.lib.CombineSequenceFileInputFormat; //导入依赖的package包/类
public static void setUpMultipleInputs(JobConf job, byte[] inputIndexes, String[] inputs, InputInfo[] inputInfos, 
		int[] brlens, int[] bclens, boolean[] distCacheOnly, boolean setConverter, ConvertTarget target) 
	throws Exception
{
	if(inputs.length!=inputInfos.length)
		throw new Exception("number of inputs and inputInfos does not match");
	
	//set up names of the input matrices and their inputformat information
	job.setStrings(INPUT_MATRICIES_DIRS_CONFIG, inputs);
	MRJobConfiguration.setMapFunctionInputMatrixIndexes(job, inputIndexes);
	
	//set up converter infos (converter determined implicitly)
	if(setConverter) {
		for(int i=0; i<inputs.length; i++)
			setInputInfo(job, inputIndexes[i], inputInfos[i], brlens[i], bclens[i], target);
	}
	
	//remove redundant inputs and pure broadcast variables
	ArrayList<Path> lpaths = new ArrayList<>();
	ArrayList<InputInfo> liinfos = new ArrayList<>();
	for(int i=0; i<inputs.length; i++)
	{
		Path p = new Path(inputs[i]);
		
		//check and skip redundant inputs
		if(   lpaths.contains(p) //path already included
		   || distCacheOnly[i] ) //input only required in dist cache
		{
			continue;
		}
		
		lpaths.add(p);
		liinfos.add(inputInfos[i]);
	}
	
	boolean combineInputFormat = false;
	if( OptimizerUtils.ALLOW_COMBINE_FILE_INPUT_FORMAT ) 
	{
		//determine total input sizes
		double totalInputSize = 0;
		for(int i=0; i<inputs.length; i++)
			totalInputSize += MapReduceTool.getFilesizeOnHDFS(new Path(inputs[i]));
			
		//set max split size (default blocksize) to 2x blocksize if (1) sort buffer large enough, 
		//(2) degree of parallelism not hurt, and only a single input (except broadcasts)
		//(the sort buffer size is relevant for pass-through of, potentially modified, inputs to the reducers)
		//(the single input constraint stems from internal runtime assumptions used to relate meta data to inputs)
		long sizeSortBuff = InfrastructureAnalyzer.getRemoteMaxMemorySortBuffer();
		long sizeHDFSBlk = InfrastructureAnalyzer.getHDFSBlockSize();
		long newSplitSize = sizeHDFSBlk * 2; //use generic config api for backwards compatibility
		double spillPercent = Double.parseDouble(job.get(MRConfigurationNames.MR_MAP_SORT_SPILL_PERCENT, "1.0"));
		int numPMap = OptimizerUtils.getNumMappers();
		if( numPMap < totalInputSize/newSplitSize && sizeSortBuff*spillPercent >= newSplitSize && lpaths.size()==1 ) {
			job.setLong(MRConfigurationNames.MR_INPUT_FILEINPUTFORMAT_SPLIT_MAXSIZE, newSplitSize);
			combineInputFormat = true;
		}
	}
	
	//add inputs to jobs input (incl input format configuration)
	for(int i=0; i<lpaths.size(); i++)
	{
		//add input to job inputs (for binaryblock we use CombineSequenceFileInputFormat to reduce task latency)
		if( combineInputFormat && liinfos.get(i) == InputInfo.BinaryBlockInputInfo )
			MultipleInputs.addInputPath(job, lpaths.get(i), CombineSequenceFileInputFormat.class);
		else
			MultipleInputs.addInputPath(job, lpaths.get(i), liinfos.get(i).inputFormatClass);
	}
}

开发者ID:apache，项目名称:systemml，代码行数:69，代码来源:MRJobConfiguration.java

注：本文中的org.apache.hadoop.mapred.lib.CombineSequenceFileInputFormat类示例整理自Github/MSDocs等源码及文档管理平台，相关代码片段筛选自各路编程大神贡献的开源项目，源码版权归原作者所有，传播和使用请参考对应项目的License；未经允许，请勿转载。

鲜花

握手

雷人

路过

鸡蛋

该文章已有0人参与评论

请发表评论

全部评论

专题导读

More+

10-27 六六分期app的软件客服如何联系？(六六分期

11-06 可心卡盟:win10系统火狐flash插件崩溃怎么

11-06 亲亲特价:怎么删除回收站图标

11-06 济南大学虚拟社区:鲁大师节能降温的具体办

11-06 xlueops.exe:无线网络安装向导

11-06 女斗合众国:win7系统cf与主机连接不稳定怎

11-06 0xc000022-[cf烟雾头]cf怎么调烟雾头

11-06 qizideyouhuo:应用程序无法正常启动0xc0000

11-06 ipz-185:win7系统vcf文件怎么打开

11-06 傻哥蹦迪:win10系统s4怎么打开usb调试

11-06 八神浩树gtaste:回收站清空了怎么恢复

11-06 妖尾之黑色守护:win10系统电脑没有1440x900

11-06 校园至尊魔王小说:win7系统浏览网页时字体

11-06 女斗合众国:win10系统访问共享文件夹提示请

11-06 tokyo hot n0654:恢复win7系统默认字体一招

11-06 雨酷仙境:设置win7系统转移临时文件夹腾出

11-06 阿穆纳伊之杖:win7系统开始菜单在右边还原

11-06 tunespotting:win10系统火狐flash插件总是

11-06 甘尔葛分析师：计谋网站seo关键词暴涨有什

11-06 蔡贵霖: 计谋网站seo关键词暴涨有什么秘密

11-06 博益网首页:ao3网页版进入不了解决方法

11-06 漏斗子专栏: 网站数据分析小白易懂精华篇

11-06 见证双虹怎么做:win7系统开启telnet命令的

11-06 颾狐蝶蜋:系统资源不足无法完成请求的服务

11-06 国光中学校歌:提交网站到alexa查询详细步骤

11-06 西安有情天:静态网页和动态网页的区别

11-06 红木雅尚斋:外部链接构造对网站的好处

11-06 前官礼遇：防止域名劫持–增强域安全性的10

11-06 密传二转答案: 中文分词算法有哪些

11-06 金泉家园邮编:百度快照劫持的表现及应对方

Java ServiceDescriptorProto类代码示例发布时间：2022-05-22

Java WebSocketClientCompressionHandler类代码示例发布时间：2022-05-22

剪的笔顺,诠释剪的笔画,认识剪的部首

1 六六分期app的软件客服如何联系？(六六分期

六六分期app的软件客服如何联系？不知道吗？加qq群【895510560】即可！标题：六六分期

阅读：19233|2023-10-27

2 可心卡盟:win10系统火狐flash插件崩溃怎么

今天小编告诉大家如何处理win10系统火狐flash插件总是崩溃的问题，可能很多用户都不知

阅读：9998|2022-11-06

3 亲亲特价:怎么删除回收站图标

今天小编告诉大家如何对win10系统删除桌面回收站图标进行设置，可能很多用户都不知道

阅读：8331|2022-11-06

4 济南大学虚拟社区:鲁大师节能降温的具体办

今天小编告诉大家如何对win10系统电脑设置节能降温的设置方法，想必大家都遇到过需要

阅读：8701|2022-11-06

5 xlueops.exe:无线网络安装向导

我们在使用xp系统的过程中,经常需要对xp系统无线网络安装向导设置进行设置，可能很多

阅读：8646|2022-11-06

6 女斗合众国:win7系统cf与主机连接不稳定怎

今天小编告诉大家如何处理win7系统玩cf老是与主机连接不稳定的问题，可能很多用户都不

阅读：9667|2022-11-06

7 0xc000022-[cf烟雾头]cf怎么调烟雾头

电脑对日常生活的重要性小编就不多说了，可是一旦碰到win7系统设置cf烟雾头的问题，很

阅读：8631|2022-11-06

8 qizideyouhuo:应用程序无法正常启动0xc0000

我们在日常使用电脑的时候，有的小伙伴们可能在打开应用的时候会遇见提示应用程序无法

阅读：8005|2022-11-06

9 ipz-185:win7系统vcf文件怎么打开

今天小编告诉大家如何对win7系统打开vcf文件进行设置，可能很多用户都不知道怎么对win

阅读：8666|2022-11-06

10 傻哥蹦迪:win10系统s4怎么打开usb调试

今天小编告诉大家如何对win10系统s4开启USB调试模式进行设置，可能很多用户都不知道怎

阅读：7540|2022-11-06

客服电话

电子邮件

Java CombineSequenceFileInputFormat类代码示例

示例1: testFormat

示例2: setUpMultipleInputs

请发表评论

全部评论

上一篇：

下一篇：

chasinginfinity/ml-from-scratch: Machine

mkyong/spring3-mvc-maven-annotation-hell

床的笔顺,关于床的笔画,体会床的部首

CVE-2022-35648

zendesk/android-floating-action-button:

剪的笔顺,诠释剪的笔画,认识剪的部首

六六分期app的软件客服如何联系？(六六分期

florent37/ViewAnimator: A fluent Android

florent37/Shrine-MaterialDesign2: implem

CVE-2020-36276

SimpleSoftwareIO/simple-sms: Send and re

关于我们

产品与服务

解决方案

139-2527-9053