https://metatext.io/datasets-list/finnish-language
Friday, 7 July 2023
List of Finnish Datasets for NLP Projects
Thursday, 29 June 2023
text summarise dataset
**Paper:**
https://arxiv.org/abs/1908.08345
**Dataset:**
1) the CNN/DailyMail news highlights dataset: somewhat Extractive
- News Articles & Related Highlights: Provides a brief overview of articles
- Input document: limited to 512 tokens
- https://www.kaggle.com/datasets/gowrishankarp/newspaper-text-summarization-cnn-dailymail
2) the New York Times Annotated Corpus (NYT): somewhat Extractive
- Contains 110,540 articles with abstract summaries
- Input document : limited to 800 tokens
- https://research.google/resources/datasets/ny-times-annotated-corpus/
3) XSum: Abstractive
- 226,711 news articles answering the question of ‘What is this articles about?’ + one-sentence summaries
- Input document: limited to 512 tokens
- https://github.com/google-research-datasets/xsum_hallucination_annotations
Popular Posts
-
This data set was created to understand the potential for machine learning, computer vision, and HPC to improve the energy efficiency aspec...
-
Best interesting data is football network refer to this page: http://www-personal.umich.edu/~mejn/netdata/
-
https://github.com/mpsilfve/finer-data The directory data contains a corpus of Finnish technology related news articles with a manually p...
-
Semi automatically generated nuclei instance segmentation and classification dataset with exhaustive nuclei labels across 19 different tiss...
-
https://github.com/RSIA-LIESMARS-WHU/RSOD-Dataset- SpaceNet[ https://spacenetchallenge.github.io/ ] https://github.com/chrieke/awesome-sa...
-
Question answering Maluuba News QA Dataset : 120K Q&A pairs on CNN news articles. Quora Question Pairs : first dataset release from ...
-
COVID-Net Open Source Initiative https://github.com/lindawangg/COVID-Net
-
Natural-Image Datasets MNIST: handwritten digits : The most commonly used sanity check. Dataset of 25x25, centered, B&W handwritten d...
-
Text Datasets 20 newsgroups : Classification task, mapping word occurences to newsgroup ID. One of the classic datasets for text classifi...