-
Arukikata Travelogue Dataset with Geographic Entity Mention, Coreference, and Link Annotation
Authors:
Shohei Higashiyama,
Hiroki Ouchi,
Hiroki Teranishi,
Hiroyuki Otomo,
Yusuke Ide,
Aitaro Yamamoto,
Hiroyuki Shindo,
Yuki Matsuda,
Shoko Wakamiya,
Naoya Inoue,
Ikuya Yamada,
Taro Watanabe
Abstract:
Geoparsing is a fundamental technique for analyzing geo-entity information in text. We focus on document-level geoparsing, which considers geographic relatedness among geo-entity mentions, and presents a Japanese travelogue dataset designed for evaluating document-level geoparsing systems. Our dataset comprises 200 travelogue documents with rich geo-entity information: 12,171 mentions, 6,339 coref…
▽ More
Geoparsing is a fundamental technique for analyzing geo-entity information in text. We focus on document-level geoparsing, which considers geographic relatedness among geo-entity mentions, and presents a Japanese travelogue dataset designed for evaluating document-level geoparsing systems. Our dataset comprises 200 travelogue documents with rich geo-entity information: 12,171 mentions, 6,339 coreference clusters, and 2,551 geo-entities linked to geo-database entries.
△ Less
Submitted 23 May, 2023;
originally announced May 2023.
-
Arukikata Travelogue Dataset
Authors:
Hiroki Ouchi,
Hiroyuki Shindo,
Shoko Wakamiya,
Yuki Matsuda,
Naoya Inoue,
Shohei Higashiyama,
Satoshi Nakamura,
Taro Watanabe
Abstract:
We have constructed Arukikata Travelogue Dataset and released it free of charge for academic research. This dataset is a Japanese text dataset with a total of over 31 million words, comprising 4,672 Japanese domestic travelogues and 9,607 overseas travelogues. Before providing our dataset, there was a scarcity of widely available travelogue data for research purposes, and each researcher had to pr…
▽ More
We have constructed Arukikata Travelogue Dataset and released it free of charge for academic research. This dataset is a Japanese text dataset with a total of over 31 million words, comprising 4,672 Japanese domestic travelogues and 9,607 overseas travelogues. Before providing our dataset, there was a scarcity of widely available travelogue data for research purposes, and each researcher had to prepare their own data. This hinders the replication of existing studies and fair comparative analysis of experimental results. Our dataset enables any researchers to conduct investigation on the same data and to ensure transparency and reproducibility in research. In this paper, we describe the academic significance, characteristics, and prospects of our dataset.
△ Less
Submitted 19 May, 2023;
originally announced May 2023.
-
Crowdsourced Hypothesis Generation and their Verification: A Case Study on Sleep Quality Improvement
Authors:
Shoko Wakamiya,
Toshiki Mera,
Eiji Aramaki,
Masaki Matsubara,
Atsuyuki Morishima
Abstract:
A clinical study is often necessary for exploring important research questions; however, this approach is sometimes time and money consuming. Another extreme approach, which is to collect and aggregate opinions from crowds, provides a result drawn from the crowds' past experiences and knowledge. To explore a solution that takes advantage of both the rigid clinical approach and the crowds' opinion-…
▽ More
A clinical study is often necessary for exploring important research questions; however, this approach is sometimes time and money consuming. Another extreme approach, which is to collect and aggregate opinions from crowds, provides a result drawn from the crowds' past experiences and knowledge. To explore a solution that takes advantage of both the rigid clinical approach and the crowds' opinion-based approach, we design a framework that exploits crowdsourcing as a part of the research process, whereby crowd workers serve as if they were a scientist conducting a "pseudo" prospective study. This study evaluates the feasibility of the proposed framework to generate hypotheses on a specified topic and verify them in the real world by employing many crowd workers. The framework comprises two phases of crowd-based workflow. In Phase 1 - the hypothesis generation and ranking phase - our system asks workers two types of questions to collect a number of hypotheses and rank them. In Phase 2 - the hypothesis verification phase - the system asks workers to verify the top-ranked hypotheses from Phase 1 by implementing one of them in real life. Through experiments, we explore the potential and limitations of the framework to generate and evaluate hypotheses about the factors that result in a good night's sleep. Our results on significant sleep quality improvement show the basic feasibility of our framework, suggesting that crowd-based research is compatible with experts' knowledge in a certain domain.
△ Less
Submitted 16 May, 2022;
originally announced May 2022.
-
Annotation-Scheme Reconstruction for "Fake News" and Japanese Fake News Dataset
Authors:
Taichi Murayama,
Shohei Hisada,
Makoto Uehara,
Shoko Wakamiya,
Eiji Aramaki
Abstract:
Fake news provokes many societal problems; therefore, there has been extensive research on fake news detection tasks to counter it. Many fake news datasets were constructed as resources to facilitate this task. Contemporary research focuses almost exclusively on the factuality aspect of the news. However, this aspect alone is insufficient to explain "fake news," which is a complex phenomenon that…
▽ More
Fake news provokes many societal problems; therefore, there has been extensive research on fake news detection tasks to counter it. Many fake news datasets were constructed as resources to facilitate this task. Contemporary research focuses almost exclusively on the factuality aspect of the news. However, this aspect alone is insufficient to explain "fake news," which is a complex phenomenon that involves a wide range of issues. To fully understand the nature of each instance of fake news, it is important to observe it from various perspectives, such as the intention of the false news disseminator, the harmfulness of the news to our society, and the target of the news. We propose a novel annotation scheme with fine-grained labeling based on detailed investigations of existing fake news datasets to capture these various aspects of fake news. Using the annotation scheme, we construct and publish the first Japanese fake news dataset. The annotation scheme is expected to provide an in-depth understanding of fake news. We plan to build datasets for both Japanese and other languages using our scheme. Our Japanese dataset is published at https://hkefka385.github.io/dataset/fakenews-japanese/.
△ Less
Submitted 6 April, 2022;
originally announced April 2022.
-
Mitigation of Diachronic Bias in Fake News Detection Dataset
Authors:
Taichi Murayama,
Shoko Wakamiya,
Eiji Aramaki
Abstract:
Fake news causes significant damage to society.To deal with these fake news, several studies on building detection models and arranging datasets have been conducted. Most of the fake news datasets depend on a specific time period. Consequently, the detection models trained on such a dataset have difficulty detecting novel fake news generated by political changes and social changes; they may possib…
▽ More
Fake news causes significant damage to society.To deal with these fake news, several studies on building detection models and arranging datasets have been conducted. Most of the fake news datasets depend on a specific time period. Consequently, the detection models trained on such a dataset have difficulty detecting novel fake news generated by political changes and social changes; they may possibly result in biased output from the input, including specific person names and organizational names. We refer to this problem as \textbf{Diachronic Bias} because it is caused by the creation date of news in each dataset. In this study, we confirm the bias, especially proper nouns including person names, from the deviation of phrase appearances in each dataset. Based on these findings, we propose masking methods using Wikidata to mitigate the influence of person names and validate whether they make fake news detection models robust through experiments with in-domain and out-of-domain data.
△ Less
Submitted 28 August, 2021;
originally announced August 2021.
-
Single Model for Influenza Forecasting of Multiple Countries by Multi-task Learning
Authors:
Taichi Murayama,
Shoko Wakamiya,
Eiji Aramaki
Abstract:
The accurate forecasting of infectious epidemic diseases such as influenza is a crucial task undertaken by medical institutions. Although numerous flu forecasting methods and models based mainly on historical flu activity data and online user-generated contents have been proposed in previous studies, no flu forecasting model targeting multiple countries using two types of data exists at present. O…
▽ More
The accurate forecasting of infectious epidemic diseases such as influenza is a crucial task undertaken by medical institutions. Although numerous flu forecasting methods and models based mainly on historical flu activity data and online user-generated contents have been proposed in previous studies, no flu forecasting model targeting multiple countries using two types of data exists at present. Our paper leverages multi-task learning to tackle the challenge of building one flu forecasting model targeting multiple countries; each country as each task. Also, to develop the flu prediction model with higher performance, we solved two issues; finding suitable search queries, which are part of the user-generated contents, and how to leverage search queries efficiently in the model creation. For the first issue, we propose the transfer approaches from English to other languages. For the second issue, we propose a novel flu forecasting model that takes advantage of search queries using an attention mechanism and extend the model to a multi-task model for multiple countries' flu forecasts. Experiments on forecasting flu epidemics in five countries demonstrate that our model significantly improved the performance by leveraging the search queries and multi-task learning compared to the baselines.
△ Less
Submitted 7 July, 2021; v1 submitted 4 July, 2021;
originally announced July 2021.
-
End-to-end Biomedical Entity Linking with Span-based Dictionary Matching
Authors:
Shogo Ujiie,
Hayate Iso,
Shuntaro Yada,
Shoko Wakamiya,
Eiji Aramaki
Abstract:
Disease name recognition and normalization, which is generally called biomedical entity linking, is a fundamental process in biomedical text mining. Recently, neural joint learning of both tasks has been proposed to utilize the mutual benefits. While this approach achieves high performance, disease concepts that do not appear in the training dataset cannot be accurately predicted. This study intro…
▽ More
Disease name recognition and normalization, which is generally called biomedical entity linking, is a fundamental process in biomedical text mining. Recently, neural joint learning of both tasks has been proposed to utilize the mutual benefits. While this approach achieves high performance, disease concepts that do not appear in the training dataset cannot be accurately predicted. This study introduces a novel end-to-end approach that combines span representations with dictionary-matching features to address this problem. Our model handles unseen concepts by referring to a dictionary while maintaining the performance of neural network-based models, in an end-to-end fashion. Experiments using two major datasets demonstrate that our model achieved competitive results with strong baselines, especially for unseen concepts during training.
△ Less
Submitted 21 April, 2021;
originally announced April 2021.
-
Influenza Surveillance using Search Engine, SNS, On-line Shopping, Q&A Service and Past Flu Patients
Authors:
Taichi Murayama,
Nobuyuki Shimizu,
Sumio Fujita,
Shoko Wakamiya,
Eiji Aramaki
Abstract:
Influenza, an infectious disease, causes many deaths worldwide. Predicting influenza victims during epidemics is an important task for clinical, hospital, and community outbreak preparation. On-line user-generated contents (UGC), primarily in the form of social media posts or search query logs, are generally used for prediction for reaction to sudden and unusual outbreaks. However, most studies re…
▽ More
Influenza, an infectious disease, causes many deaths worldwide. Predicting influenza victims during epidemics is an important task for clinical, hospital, and community outbreak preparation. On-line user-generated contents (UGC), primarily in the form of social media posts or search query logs, are generally used for prediction for reaction to sudden and unusual outbreaks. However, most studies rely only on the UGC as their resource and do not use various UGCs. Our study aims to solve these questions about Influenza prediction: Which model is the best? What combination of multiple UGCs works well? What is the nature of each UGC? We adapt some models, LASSO Regression, Huber Regression, Support Vector Machine regression with Linear kernel (SVR) and Random Forest, to test the influenza volume prediction in Japan during 2015 - 2018. For that, we use on-line five data resources: (1) past flu patients, (2) SNS (Twitter), (3) search engines (Yahoo! Japan), (4) shopping services (Yahoo! Shopping), and (5) Q&A services (Yahoo! Chiebukuro) as resources of each model. We then validate respective resources contributions using the best model, Huber Regression, with all resources except one resource. Finally, we use Bayesian change point method for ascertaining whether the trend of time series on any resources is reflected in the trend of flu patient count or not. Our experiments show Huber Regression model based on various data resources produces the most accurate results. Then, from the change point analysis, we get the result that search query logs and social media posts for three years represents these resources as a good predictor. Conclusions: We show that Huber Regression based on various data resources is strong for outliers and is suitable for the flu prediction. Additionally, we indicate the characteristics of each resource for the flu prediction.
△ Less
Submitted 14 April, 2021;
originally announced April 2021.
-
KART: Parameterization of Privacy Leakage Scenarios from Pre-trained Language Models
Authors:
Yuta Nakamura,
Shouhei Hanaoka,
Yukihiro Nomura,
Naoto Hayashi,
Osamu Abe,
Shuntaro Yada,
Shoko Wakamiya,
Eiji Aramaki
Abstract:
For the safe sharing pre-trained language models, no guidelines exist at present owing to the difficulty in estimating the upper bound of the risk of privacy leakage. One problem is that previous studies have assessed the risk for different real-world privacy leakage scenarios and attack methods, which reduces the portability of the findings. To tackle this problem, we represent complex real-world…
▽ More
For the safe sharing pre-trained language models, no guidelines exist at present owing to the difficulty in estimating the upper bound of the risk of privacy leakage. One problem is that previous studies have assessed the risk for different real-world privacy leakage scenarios and attack methods, which reduces the portability of the findings. To tackle this problem, we represent complex real-world privacy leakage scenarios under a universal parameterization, \textit{Knowledge, Anonymization, Resource, and Target} (KART). KART parameterization has two merits: (i) it clarifies the definition of privacy leakage in each experiment and (ii) it improves the comparability of the findings of risk assessments. We show that previous studies can be simply reviewed by parameterizing the scenarios with KART. We also demonstrate privacy risk assessments in different scenarios under the same attack method, which suggests that KART helps approximate the upper bound of risk under a specific attack or scenario. We believe that KART helps integrate past and future findings on privacy risk and will contribute to a standard for sharing language models.
△ Less
Submitted 17 March, 2022; v1 submitted 31 December, 2020;
originally announced January 2021.
-
Universal Fake News Collection System using Debunking Tweets
Authors:
Taichi Murayama,
Shoko Wakamiya,
Eiji Aramaki
Abstract:
Large numbers of people use Social Networking Services (SNS) for easy access to various news, but they have more opportunities to obtain and share ``fake news'' carrying false information. Partially to combat fake news, several fact-checking sites such as Snopes and PolitiFact have been founded. Nevertheless, these sites rely on time-consuming and labor-intensive tasks. Moreover, their available l…
▽ More
Large numbers of people use Social Networking Services (SNS) for easy access to various news, but they have more opportunities to obtain and share ``fake news'' carrying false information. Partially to combat fake news, several fact-checking sites such as Snopes and PolitiFact have been founded. Nevertheless, these sites rely on time-consuming and labor-intensive tasks. Moreover, their available languages are not extensive. To address these difficulties, we propose a new fake news collection system based on rule-based (unsupervised) frameworks that can be extended easily for various languages. The system collects news with high probability of being fake by debunking tweets by users and presents event clusters gathering higher attention. Our system currently functions in two languages: English and Japanese. It shows event clusters, 65\% of which are actually fake. In future studies, it will be applied to other languages and will be published with a large fake news dataset.
△ Less
Submitted 28 July, 2020;
originally announced July 2020.
-
Modeling the spread of fake news on Twitter
Authors:
Taichi Murayama,
Shoko Wakamiya,
Eiji Aramaki,
Ryota Kobayashi
Abstract:
Fake news can have a significant negative impact on society because of the growing use of mobile devices and the worldwide increase in Internet access. It is therefore essential to develop a simple mathematical model to understand the online dissemination of fake news. In this study, we propose a point process model of the spread of fake news on Twitter. The proposed model describes the spread of…
▽ More
Fake news can have a significant negative impact on society because of the growing use of mobile devices and the worldwide increase in Internet access. It is therefore essential to develop a simple mathematical model to understand the online dissemination of fake news. In this study, we propose a point process model of the spread of fake news on Twitter. The proposed model describes the spread of a fake news item as a two-stage process: initially, fake news spreads as a piece of ordinary news; then, when most users start recognizing the falsity of the news item, that itself spreads as another news story. We validate this model using two datasets of fake news items spread on Twitter. We show that the proposed model is superior to the current state-of-the-art methods in accurately predicting the evolution of the spread of a fake news item. Moreover, a text analysis suggests that our model appropriately infers the correction time, i.e., the moment when Twitter users start realizing the falsity of the news item. The proposed model contributes to understanding the dynamics of the spread of fake news on social media. Its ability to extract a compact representation of the spreading pattern could be useful in the detection and mitigation of fake news.
△ Less
Submitted 27 April, 2021; v1 submitted 28 July, 2020;
originally announced July 2020.
-
Fake News Detection using Temporal Features Extracted via Point Process
Authors:
Taichi Murayama,
Shoko Wakamiya,
Eiji Aramaki
Abstract:
Many people use social networking services (SNSs) to easily access various news. There are numerous ways to obtain and share ``fake news,'' which are news carrying false information. To address fake news, several studies have been conducted for detecting fake news by using SNS-extracted features. In this study, we attempt to use temporal features generated from SNS posts by using a point process a…
▽ More
Many people use social networking services (SNSs) to easily access various news. There are numerous ways to obtain and share ``fake news,'' which are news carrying false information. To address fake news, several studies have been conducted for detecting fake news by using SNS-extracted features. In this study, we attempt to use temporal features generated from SNS posts by using a point process algorithm to identify fake news from real news. Temporal features in fake news detection have the advantage of robustness over existing features because it has minimal dependence on fake news propagators. Further, we propose a novel multi-modal attention-based method, which includes linguistic and user features alongside temporal features, for detecting fake news from SNS posts. Results obtained from three public datasets indicate that the proposed model achieves better performance compared to existing methods and demonstrate the effectiveness of temporal features for fake news detection.
△ Less
Submitted 28 July, 2020;
originally announced July 2020.
-
Syndromic surveillance using search query logs and user location information from smartphones against COVID-19 clusters in Japan
Authors:
Shohei Hisada,
Taichi Murayama,
Kota Tsubouchi,
Sumio Fujita,
Shuntaro Yada,
Shoko Wakamiya,
Eiji Aramaki
Abstract:
[Background] Two clusters of coronavirus disease 2019 (COVID-19) were confirmed in Hokkaido, Japan in February 2020. To capture the clusters, this study employs Web search query logs and user location information from smartphones. [Material and Methods] First, we anonymously identified smartphone users who used a Web search engine (Yahoo! JAPAN Search) for the COVID-19 or its symptoms via its comp…
▽ More
[Background] Two clusters of coronavirus disease 2019 (COVID-19) were confirmed in Hokkaido, Japan in February 2020. To capture the clusters, this study employs Web search query logs and user location information from smartphones. [Material and Methods] First, we anonymously identified smartphone users who used a Web search engine (Yahoo! JAPAN Search) for the COVID-19 or its symptoms via its companion application for smartphones (Yahoo Japan App). We regard these searchers as Web searchers who are suspicious of their own COVID-19 infection (WSSCI). Second, we extracted the location of the WSSCI via the smartphone application. The spatio-temporal distribution of the number of WSSCI are compared with the actual location of the known two clusters. [Result and Discussion] Before the early stage of the cluster development, we could confirm several WSSCI, which demonstrated the basic feasibility of our WSSCI-based approach. However, it is accurate only in the early stage, and it was biased after the public announcement of the cluster development. For the case where the other cluster-related resources, such as fine-grained population statistics, are not available, the proposed metric would be helpful to catch the hint of emerging clusters.
△ Less
Submitted 21 April, 2020;
originally announced April 2020.
-
NAIST COVID: Multilingual COVID-19 Twitter and Weibo Dataset
Authors:
Zhiwei Gao,
Shuntaro Yada,
Shoko Wakamiya,
Eiji Aramaki
Abstract:
Since the outbreak of coronavirus disease 2019 (COVID-19) in the late 2019, it has affected over 200 countries and billions of people worldwide. This has affected the social life of people owing to enforcements, such as "social distancing" and "stay at home." This has resulted in an increasing interaction through social media. Given that social media can bring us valuable information about COVID-1…
▽ More
Since the outbreak of coronavirus disease 2019 (COVID-19) in the late 2019, it has affected over 200 countries and billions of people worldwide. This has affected the social life of people owing to enforcements, such as "social distancing" and "stay at home." This has resulted in an increasing interaction through social media. Given that social media can bring us valuable information about COVID-19 at a global scale, it is important to share the data and encourage social media studies against COVID-19 or other infectious diseases. Therefore, we have released a multilingual dataset of social media posts related to COVID-19, consisting of microblogs in English and Japanese from Twitter and those in Chinese from Weibo. The data cover microblogs from January 20, 2020, to March 24, 2020. This paper also provides a quantitative as well as qualitative analysis of these datasets by creating daily word clouds as an example of text-mining analysis. The dataset is now available on Github. This dataset can be analyzed in a multitude of ways and is expected to help in efficient communication of precautions related to COVID-19.
△ Less
Submitted 17 April, 2020;
originally announced April 2020.
-
Density Estimation for Geolocation via Convolutional Mixture Density Network
Authors:
Hayate Iso,
Shoko Wakamiya,
Eiji Aramaki
Abstract:
Nowadays, geographic information related to Twitter is crucially important for fine-grained applications. However, the amount of geographic information avail- able on Twitter is low, which makes the pursuit of many applications challenging. Under such circumstances, estimating the location of a tweet is an important goal of the study. Unlike most previous studies that estimate the pre-defined dist…
▽ More
Nowadays, geographic information related to Twitter is crucially important for fine-grained applications. However, the amount of geographic information avail- able on Twitter is low, which makes the pursuit of many applications challenging. Under such circumstances, estimating the location of a tweet is an important goal of the study. Unlike most previous studies that estimate the pre-defined district as the classification task, this study employs a probability distribution to represent richer information of the tweet, not only the location but also its ambiguity. To realize this modeling, we propose the convolutional mixture density network (CMDN), which uses text data to estimate the mixture model parameters. Experimentally obtained results reveal that CMDN achieved the highest prediction performance among the method for predicting the exact coordinates. It also provides a quantitative representation of the location ambiguity for each tweet that properly works for extracting the reliable location estimations.
△ Less
Submitted 8 May, 2017;
originally announced May 2017.