2020 | 19 | 18 | 17 | 16 | 15 | 14 | 13 | 12 | 11 | 10 | 09 | 08 | 07 | 06 | 05 | 04 | 03 | 02 | 01 | 00 | 99

Data Mining Research

SourcererCC: Scaling Type-3 Clone Detection to Large Software Repositories

Given the availability of large-scale source-code repositories, there have been a large number of applications for clone detection. Unfortunately, despite a decade of active research, there is a marked lack in clone detectors that scale to large software repositories. In particular for detecting near-miss clones where significant editing activities may take place in the cloned code.

Principal Investigators:

Cristina Videira Lopes

Research Area(s):

Large Systems and Large Data,

Clone Detection,

Data Mining

Project Dates:

January 2014

Scaling Token-Based Code Clone Detection

We developed a token-based approach for large scale code clone detection which is based on a filtering heuristic that reduces the number of token comparisons when the two code blocks are compared. We also developed a MapReduce based parallel algorithm that uses the filtering heuristic and scales to thousands of projects. The filtering heuristic is generic and can also be used in conjunction with other token-based approaches. In that context, we demonstrated how it can increase the retrieval speed and decrease the memory usage of the index-based approaches.

Principal Investigators:

Cristina Videira Lopes

Research Area(s):

Large Systems and Large Data,

Clone Detection,

Data Mining

Project Dates:

July 2011 to January 2014

The Sourcerer Project

Sourcerer is an ongoing research project at the University of California, Irvine aimed at exploring open source projects through the use of code analysis. The existence of an extremely large body of open source code presents a tremendous opportunity for software engineering research. Not only do we leverage this code for our own research, but we provide the open source Sourcerer Infrastructure and curated datasets for other researchers to use.

The Sourcerer Infrastructure is composed of a number of layers.

Principal Investigators:

Cristina Videira Lopes

Research Area(s):

Large Systems and Large Data,

Open Source,

Data Mining,

Software Analysis

Project Dates:

January 2006