Academic staff

Stamatatos Efstathios

Personal Information
Stamatatos Efstathios

Professor


stamatatos [at] aegean [dot] gr

+30.22730.82260

Lyberi building, A7

Thursday 15:00-17:00

Personal Website

Citations (Google Scholar)

Copyright Notice: This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. In most cases, these works may not be reposted or mass reproduced without the explicit permission of the copyright holder.


Journal Publications

[1]
N. Manousakis, E. Stamatatos, Authorship Analysis and the Ending of Seven Against Thebes: Aeschylus' Antigone or Updating Adaptation?, Classical World, Vol. 116, No. 3, pp. 247-274, 2023, John Hopkins University Press

[1]
G. Barlas, E. Stamatatos, A Transfer Learning Approach to Cross domain Authorship Attribution, Evolving Systems, Vol. 12, pp. 625-643, 2021, Springer, https://link.springer.com/article/10.100...

[1]
F. Sánchez-Vega, E. Villatoro-Tello, M. Montes-y-Gómez, P. Rosso, E. Stamatatos, L. Villaseñor-Pineda, Paraphrase Plagiarism Identification with Character-level Features, Pattern Analysis and Applications, Vol. 22, No. 2, pp. 669-681, 2019, Springer, http://dx.doi.org/10.1007/s10044-017-067...
[2]
N. Potha, E. Stamatatos, Improving author verification based on topic modeling, Journal of the Association for Information Science and Technology , 2019, Wiley Online Library

[1]
D. Pritsos, E. Stamatatos, Open Set Evaluation in Web Genre Identification, Language Resources and Evaluation, Vol. 52, No. 4, pp. 949–968, 2018, Springer, http://dx.doi.org/10.1007/s10579-018-941...
[2]
E. Stamatatos, Masking Topic-related Information to Enhance Authorship Attribution, Journal of the Association for Information Science and Technology, Vol. 69, No. 3, pp. 461-473, 2018, Wiley, https://doi.org/10.1002/asi.23968

[1]
S. Miranda-Jiménez, E. Stamatatos, Automatic Generation of Summary Obfuscation Corpus for Plagiarism Detection, Acta Polytechnica Hungarica, Vol. 14, No. 3, pp. 99-112, 2017, Budapest Tech Polytechnical Institution, https://www.uni-obuda.hu/journal/Miranda...
[2]
N. Manousakis, E. Stamatatos, Devising Rhesus: A strange ‘collaboration’ between Aeschylus and Euripides, Digital Scholarship in the Humanities, 2017, (to_appear), https://doi.org/10.1093/llc/fqx021
[3]
A. Rocha, W.J. Scheirer, C.W. Forstall, T. Cavalcante, A. Theophilo, B. Shen, A.R.B. Carvalho, E. Stamatatos, Authorship Attribution for Social Media Forensics, IEEE Transactions on Information Forensics and Security, Vol. 12, No. 1, pp. 5-33, 2017, IEEE, https://doi.org/10.1109/TIFS.2016.260396..., indexed in SCI-E

[1]
A. Pastor López-Monroy, M. Montes-y-Gómez, H.J. Escalante, L. Villaseñor-Pineda, E. Stamatatos, Discriminative Subprofile-Specific Representations for Author Profiling in Social Media, Knowledge-Based Systems, Vol. 89, pp. 134–147, 2015, Elsevier, http://dx.doi.org/10.1016/j.knosys.2015...., indexed in SCI-E

[1]
G. Sidorov, F. Velasquez, E. Stamatatos, A. Gelbukh, L. Chanona-Hernández, Syntactic N-grams as Machine Learning Features for Natural Language Processing, Expert Systems with Applications, Vol. 41, No. 3, pp. 853-860, 2014, Elsevier, http://dx.doi.org/10.1016/j.eswa.2013.08..., indexed in SCI-E

G. Sidorov, F. Velasquez, E. Stamatatos, A. Gelbukh, L. Chanona-Hernández, Syntactic N-grams as Machine Learning Features for Natural Language Processing, Expert Systems with Applications, 2013, (to_appear), http://dx.doi.org/10.1016/j.eswa.2013.08...
Abstract:
In this paper we introduce and discuss a concept of syntactic n-grams (sn-grams). Sn-grams differ from traditional n-grams in the manner how we construct them, i.e., what elements are considered neighbors. In case of sngrams, the neighbors are taken by following syntactic relations in syntactic trees, and not by taking words as they appear in a text, i.e., sn-grams are constructed by following paths in syntactic trees. In this manner, sn-grams allow bringing syntactic knowledge into machine learning methods; still, previous parsing is necessary for their construction. Sn-grams can be applied in any natural language processing (NLP) task where traditional n-grams are used. We describe how sn-grams were applied to authorship attribution. We used as baseline traditional n-grams of words, part of speech (POS) tags and characters; three classifiers were applied: support vector machines (SVM), naive Bayes (NB), and tree classifier J48. Sn-grams give better results with SVM classifier.
E. Stamatatos, On the Robustness of Authorship Attribution Based on Character n-gram Features, Journal of Law and Policy, Vol. 21, No. 2, pp. 421-439, 2013, Brooklyn Law School, http://practicum.brooklaw.edu/journals/j...
Abstract:
A number of independent authorship attribution studies have demonstrated the effectiveness of character n-gram features for representing the stylistic properties of text. However, the vast majority of these studies examined the simple case where the training and test corpora are similar in terms of genre, topic, and distribution of the texts. Hence, there are doubts whether such a simple and low-level representation is equally effective in realistic conditions where some of the above factors are not possible to remain stable. In this study, the robustness of authorship attribution based on character n-gram features is tested under cross-genre and cross-topic conditions. In addition, the distribution of texts over the candidate authors varies in training and test corpora to imitate real cases. Comparative results with another competitive text representation approach based on very frequent words show that character n-grams are better able to capture stylistic properties of text when there are significant differences among the training and test corpora. Moreover, a set of guidelines to tune an authorship attribution model according to the properties of training and test corpora is provided.

E. Stamatatos, Plagiarism Detection Using Stopword n-grams, Journal of the American Society for Information Science and Technology, Vol. 62, No. 12, pp. 2512-2527, 2011, Wiley, http://dx.doi.org/10.1002/asi.21630, indexed in SCI-E
Abstract:
In this paper, a novel method for detecting plagiarized passages in document collections is presented. In contrast to previous work in this field that uses content terms to represent documents, the proposed method is based on a small list of stopwords (i.e., very frequent words). We show that stopword n-grams reveal important information for plagiarism detection since they are able to capture syntactic similarities between suspicious and original documents and they can be used to detect the exact plagiarized passage boundaries. Experimental results on a publicly-available corpus demonstrate that the performance of the proposed approach is competitive when compared with the best reported results. More importantly, it achieves significantly better results when dealing with difficult plagiarism cases where the plagiarized passages are highly modified and most of the words or phrases have been replaced with synonyms.
G. Frantzeskou, S. MacDonell, E. Stamatatos, S. Georgiou, S. Gritzalis, The significance of user-defined identifiers in Java source code authorship identification, International Journal of Computer Systems Science and Engineering, Vol. 26, No. 2, pp. 139-148, 2011, http://aut.researchgateway.ac.nz/bitstre..., indexed in SCI-E, IF = 0.371
Abstract:
When writing source code, programmers have varying levels of freedom when it comes to the creation and use of identifiers. Do they habitually use the same identifiers, names that are different to those used by others? Is it then possible to tell who the author of a piece of code is by examining these identifiers?

I. Kanaris, E. Stamatatos, Learning to Recognize Webpage Genres, Information Processing and Management, Vol. 45, No. 5, pp. 499-512, 2009, Elsevier, http://dx.doi.org/10.1016/j.ipm.2009.05...., indexed in SCI-E
Abstract:
Webpages are mainly distinguished by their topic (e.g., politics, sports etc.) and genre (e.g., blogs, homepages, e-shops, etc.). Automatic detection of webpage genre could considerably enhance the ability of modern search engines to focus on the requirements of the userメs information need. In this paper, we present an approach to webpage genre detection based on a fully-automated extraction of the feature set that represents the style of webpages. The features we propose (character n-grams of variable length and HTML tags) are language- independent and easily-extracted while they can be adapted to the properties of the still evolving web genres and the noisy environment of the web. Experiments based on two publicly-available corpora show that the performance of the proposed approach is superior in comparison to previously reported results. It is also shown that character n-grams are better features than words when the dimensionality increases while the binary representation is more effective than the term-frequency representation for both feature types. Moreover, we perform a series of cross-check experiments (e.g., training using a genre palette and testing using a different genre palette as well as using the features extracted from one corpus to discriminate the genres of the other corpus) to illustrate the robustness of our approach and its ability to capture the general stylistic properties of genre categories even when the feature set is not optimized for the given corpus.
E. Stamatatos, A Survey of Modern Authorship Attribution Methods, Journal of the American Society for Information Science and Technology, Vol. 60, No. 3, pp. 538-556, 2009, Wiley, http://dx.doi.org/10.1002/asi.21001
Abstract:
Authorship attribution supported by statistical or computational methods has a long history starting from the 19th century and is marked by the seminal study of Mosteller and Wallace (1964) on the authorship of the disputed “Federalist Papers.”During the last decade, this scientific field has been developed substantially, taking advantage of research advances in areas such as machine learning, information retrieval, and natural language processing. The plethora of available electronic texts (e.g., e-mail messages, online forum messages, blogs, source code, etc.) indicates a wide variety of applications of this technology, provided it is able to handle short and noisy text from multiple candidate authors. In this article, a survey of recent advances of the automated approaches to attributing authorship is presented, examining their characteristics for both text representation and text classification. The focus of this survey is on computational requirements and settings rather than on linguistic or literary issues. We also discuss evaluation methodologies and criteria for authorship attribution studies and list open questions that will attract future work in this area.

E. Stamatatos, Author Identification: Using Text Sampling to Handle the Class Imbalance Problem, Information Processing and Management, Vol. 44, No. 2, pp. 790-799, 2008, Elsevier, http://dx.doi.org/10.1016/j.ipm.2007.05....
Abstract:
Authorship analysis of electronic texts assists digital forensics and anti-terror investigation. Author identification can be seen as a single-label multi-class text categorization problem. Very often, there are extremely few training texts at least for some of the candidate authors or there is a significant variation in the text-length among the available training texts of the candidate authors. Moreover, in this task usually there is no similarity between the distribution of training and test texts over the classes, that is, a basic assumption of inductive learning does not apply. In this paper, we present methods to handle imbalanced multi-class textual datasets. The main idea is to segment the training texts into text samples according to the size of the class, thus producing a fairer classification model. Hence, minority classes can be segmented into many short samples and majority classes into less and longer samples. We explore text sampling methods in order to construct a training set according to a desirable distribution over the classes. Essentially, by text sampling we provide new synthetic data that artificially increase the training size of a class. Based on two text corpora of two languages, namely, newswire stories in English and newspaper reportage in Arabic, we present a series of authorship identification experiments on various multiclass imbalanced cases that reveal the properties of the presented methods.
G. Frantzeskou, S. MacDonell, E. Stamatatos, S. Gritzalis, Examining the Significance of high-level programming features in Source-code Author Classification, Journal of Systems and Software, Vol. 81, No. 3, pp. 447-460, 2008, Elsevier, http://www.sciencedirect.com/science/art..., indexed in SCI-E, IF = 1.241
Abstract:
The use of Source Code Author Profiles (SCAP) represents a new, highly accurate approach to source code authorship identification that is, unlike previous methods, language independent. While accuracy is clearly a crucial requirement of any author identification method, in cases of litigation regarding authorship, plagiarism, and so on, there is also a need to know why it is claimed that a piece of code is written by a particular author. What is it about that piece of code that suggests a particular author? What features in the code make one author more likely than another? In this study, we describe a means of identifying the high-level features that contribute to source code authorship identification using as a tool the SCAP method. A variety of features are considered for Java and Common Lisp and the importance of each feature in determining authorship is measured through a sequence of experiments in which we remove one feature at a time. The results show that, for these programs, comments, layout features and package-related naming influence classification accuracy whereas user-defined naming, an obvious programmer related feature, does not appear to influence accuracy. A comparison is also made between the relative feature contributions in programs written in the two languages.

G. Frantzeskou, E. Stamatatos, S. Gritzalis, C. Chaski, B. Howald, Identifying Authorship by Byte Level n-grams: The Source Code Author Profile (SCAP) Method, International Journal of Digital Evidence, Vol. 6, No. 1, pp. 1-15, 2007, Economic Crime Institute, http://www.utica.edu/academic/institutes...
Abstract:
Source code author identification deals with identifying the most likely author of a computer program, given a set of predefined author candidates. There are several scenarios where digital evidence of this kind plays a role in investigation and adjudication, such as code authorship disputes, intellectual property infringement, tracing the source of code left in the system after a cyber attack, and so forth. As in any identification task, the disputed program is compared to undisputed, known programming samples by the predefined author candidates. We present a new approach, called the SCAP (Source Code Author Profiles) approach, based on byte-level n-gram profiles representing the source code author’s style. The SCAP method extends a method originally applied to natural language text authorship attribution; we show that an n-gram approach also suits the characteristics of source code analysis. The methodological extension includes a simplified profile and a less complicated, but more effective, similarity measure. Experiments on data sets of different programming-language (Java or C++) and commented/commentless code demonstrate the effectiveness of these extensions. The SCAP approach is programming-language independent. Moreover, the SCAP approach deals surprisingly well with cases where only a limited amount of very short programs per programmer is available for training. Finally, it is also demonstrated that SCAP effectiveness persists even in the absence of comments in the source code, a condition usually met in cyber-crime cases.
B. Stein, S.Argamon, E. Stamatatos, Plagiarism Analysis, Authorship Identification, and Near-Duplicate Detection, ACM SIGIR Forum, Vol. 41, No. 2, pp. 68-71, 2007, http://dx.doi.org/10.1145/1328964.132897...
Abstract:
Goal of the workshop was to bring together experts and prospective researchers around the exciting and future-oriented topic of plagiarism analysis, authorship identi¯cation, and high similarity search. This topic receives increasing attention, which results, among others, from the fact that information about nearly any subject can be found on the World Wide Web.
I. Kanaris, K. Kanaris, J. Houvardas, E. Stamatatos, Words vs. Character N-grams for Anti-spam Filtering, Int. Journal on Artificial Intelligence Tools, Vol. 16, No. 6, pp. 1047-1067, 2007, World Scientific, http://dx.doi.org/10.1142/S0218213007003...
Abstract:
The increasing number of unsolicited e-mail messages (spam) reveals the need for the development of reliable anti-spam filters. The vast majority of content-based techniques rely on word-based representation of messages. Such approaches require reliable tokenizers for detecting the token boundaries. As a consequence, a common practice of spammers is to attempt to confuse tokenizers using unexpected punctuation marks or special characters within the message. In this paper we explore an alternative low-level representation based on character n-grams which avoids the use of tokenizers and other language-dependent tools. Based on experiments on two well-known benchmark corpora and a variety of evaluation measures, we show that character n-grams are more reliable features than word-tokens despite the fact that they increase the dimensionality of the problem. Moreover, we propose a method for extracting variable-length n-grams which produces optimal classifiers among the examined models under cost-sensitive evaluation.

E. Stamatatos, Authorship Attribution Based on Feature Set Subspacing Ensembles, Int. Journal on Artificial Intelligence Tools, Vol. 15, No. 5, pp. 823-838, 2006, World Scientific, http://dx.doi.org/10.1142/S0218213007003...
Abstract:
Authorship attribution can assist the criminal investigation procedure as well as cybercrime analysis. This task can be viewed as a single-label multi-class text categorization problem. Given that the style of a text can be represented as mere word frequencies selected in a language-independent method, suitable machine learning techniques able to deal with high dimensional feature spaces and sparse data can be directly applied to solve this problem. This paper focuses on classifier ensembles based on feature set subspacing. It is shown that an effective ensemble can be constructed using, exhaustive disjoint subspacing, a simple method producing many poor but diverse base classifiers. The simple model can be enhanced by a variation of the technique of cross-validated committees applied to the feature set. Experiments on two benchmark text corpora demonstrate the effectiveness of the presented method improving previously reported results and compare it to support vector machines, an alternative suitable machine learning approach to authorship attribution.

E. Stamatatos, G. Widmer, Automatic Identification of Music Performers with Learning Ensembles, Artificial Intelligence, Vol. 165, No. 1, pp. 37-56, 2005, Elsevier, http://dx.doi.org/10.1016/j.artint.2005....
Abstract:
This paper addresses the problem of identifying the most likely music performer, given a set of performances of the same piece by a number of skilled candidate pianists. We propose a set of features for representing the stylistic characteristics of a music performer, introducing norm-based features that are relevant to the average performance. A database of piano performances of 22 pianists playing two pieces by F. Chopin is used in the presented experiments. Due to the limitations of the training set size and the characteristics of the input features we propose an ensemble of simple classifiers derived by both subsampling the training set and subsampling the input features. The presented experiments show that the proposed features are able to quantify the differences between music performers. The proposed ensemble can efficiently cope with multi-class music performer recognition under inter-piece conditions, a difficult musical task, displaying a level of accuracy unlikely to be matched by human listeners (under similar conditions). Moreover, it is empirically demonstrated that the average performance is at least as effective as the best of the constituent individual performances while ‘extreme’ performances have the lowest discriminatory potential when used as norm.

E. Stamatatos, N. Fakotakis, G. Kokkinakis, Computer-Based Authorship Attribution without Lexical Measures, Computers and the Humanities, Vol. 35, No. 2, pp. 193-214, 2001, Kluwer, http://dx.doi.org/10.1023/A:100268191951...
Abstract:
The most important approaches to computer-assisted authorship attribution are exclusively based on lexical measures that either represent the vocabulary richness of the author or simply comprise frequencies of occurrence of common words. In this paper we present a fully-automated approach to the identification of the authorship of unrestricted text that excludes any lexical measure. Instead we adapt a set of style markers to the analysis of the text performed by an already existing natural language processing tool using three stylometric levels, i.e., token-level, phrase-level, and analysis-level measures. The latter represent the way in which the text has been analyzed. The presented experiments on a Modern Greek newspaper corpus show that the proposed set of style markers is able to distinguish reliably the authors of a randomly-chosen group and performs better than a lexically-based approach. However, the combination of these two approaches provides the most accurate solution (i.e., 87% accuracy). Moreover, we describe experiments on various sizes of the training data as well as tests dealing with the significance of the proposed set of style markers.

E. Stamatatos, N. Fakotakis, G. Kokkinakis, Automatic Text Categorization in Terms of Genre and Author, Computational Linguistics, Vol. 26, No. 4, pp. 461-485, 2000, MIT Press, http://dx.doi.org/10.1162/08912010075010...
Abstract:
The two main factors that characterize a text are its content and its style, and both can be used as a means of categorization. In this paper we present an approach to text categorization in terms of genre and author for Modern Greek. In contrast to previous stylometric approaches, we attempt to take full advantage of existing natural language processing (NLP) tools. To this end, we propose a set of style markers including analysis-levelmeasures that represent the way in which the input text has been analyzed and capture useful stylistic information without additional cost. We present a set of small-scale but reasonable experiments in text genre detection, author identiŽcation, and author veriŽcation tasks and show that the proposed method performs better than the most popular distributional lexical measures, i.e., functions of vocabulary richness and frequencies of occurrence of the most frequent words. All the presented experiments are based on unrestricted text downloaded from the World Wide Web without any manual text preprocessing or text sampling.Various performance issues regarding the training set size and the signiŽcance of the proposed style markers are discussed.Our system can be used in any application that requires fast and easily adaptable text categorization in terms of stylistically homogeneous categories. Moreover, the procedure of deŽning analysis-level markers can be followed in order to extract useful stylistic information using existing text processing tools.

S. Michos, E. Stamatatos, N. Fakotakis, Supporting Multilinguality in Library Automation Systems Using AI Tools, Applied Artificial Intelligence, Vol. 13, No. 7, pp. 679-704, 1999, Taylor & Francis, http://dx.doi.org/10.1080/08839519911724...
Abstract:
Language barriers present a major problemin the e€ ectiveness of resource sharing and in common access to the resources of libraries. In this paper we present the TRANSLIB system, which consists of an integration of both new and existing multilingual information tools. This systemtakes full advantage of some AI± based methods in order to provide multilingual access to library catalogues. Its main features include functionalities for searching in multiple languages, multilingual presentation of the query results, and localization of the user interface. TRANSLIB has currently been tested in existing medium± sized bibliographic databases. Evaluation results show a remarkable improvement in the search process and report high user friendliness and easy and low± cost maintenance and upgrade of the system.
Contact
  • President: Skoutas Dimitrios
  • Secretariat Head: Karagianni Kalliopi
  • Undergraduate Secretariat: ICS Eng. Department
  • Postgraduate Secretariat: ICS Eng. Department
  • Email: dicsd [at] aegean [dot] gr
  • Phone: 2273082000
  • Address: Κτήριο Λυμπέρη, Παλαμά 2 & Γοργύρας, Τ.Κ. 83200
  • Website: www.icsd.aegean.gr
  • Office Hours: Δευτέρα - Παρασκευή: 8:00 - 16:00
Στατιστικά Σπουδών
Μέσος Όρος Βαθμού Πτυχίου

7.76

Μέσος χρόνος Απόκτησης Πτυχίου

6.5 έτη

Μαθήματα με εργαστήριο

46

Κύκλοι Σπουδών

6

Μαθήματα Υποχρεωτικά

36

Μαθήματα Κύκλου

8

Σύνολο μαθημάτων για πτυχίο

55

Διπλωματική Εργασία

Υποχρεωτική