基于网页结构相似度的Web 信息抽取 Information Extraction Based on Structural Similarities between Web Pages期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

按检索

基于网页结构相似度的Web 信息抽取

引用本文：	聂卉.基于网页结构相似度的Web 信息抽取[J].情报学报,2011,28(3).

作者姓名：	聂卉

作者单位：	中山大学资讯管理系,广州,510275

基金项目：	教育部人文社会科学研究项目

摘要：	本文重点探讨基于编辑距离的网页相似度算法在Web 抽取系统中的应用与实现.通过结合基于URL 及编辑距离的网页结构相似度的计算方法,抽取系统在抽取过程中能够检测网页结构的变化,从而主动做出判断,选择适应规则进行抽取或通过主动学习自动扩展规则库.结构相似度计算赋予系统感知网页结构变化的能力,系统通过主动自我更新与调整,能更好地适应面向实际应用的异构资源的获取.算法的可行性和效率在原型系统中得以验证.
关键词：	信息抽取结构相似度编辑距离
Information Extraction Based on Structural Similarities between Web Pages

Nie Hui.Information Extraction Based on Structural Similarities between Web Pages[J].Journal of the China Society for Scientific andTechnical Information,2011,28(3).

Authors:	Nie Hui

Institution:	Nie Hui (School of Information Management,Sun Yat-Sen University,Guangzhou 510275)

Abstract:	In this paper,we focus on Web Information Extraction by use of structural similarities.Combined with URL similarity,the Tree- Edit- Distance method is adapted to measure the similarity between web pages.The similarity is used for detecting changes between the pages as the criterion.By this way,the suitable extraction rules will be selected,otherwise, the target page will be fed to rule learner module.Extraction rules are induced automatically by machine learning.The introduction of the algorithm given extra...

Keywords:	Web
本文献已被 CNKI 万方数据等数据库收录！

设为首页 | 免责声明 | 关于勤云 | 加入收藏