# 哈佛把十億篇美國報紙文章通通電子化啦！

**URL:** https://vip.studycamp.tw/t/%E5%93%88%E4%BD%9B%E6%8A%8A%E5%8D%81%E5%84%84%E7%AF%87%E7%BE%8E%E5%9C%8B%E5%A0%B1%E7%B4%99%E6%96%87%E7%AB%A0%E9%80%9A%E9%80%9A%E9%9B%BB%E5%AD%90%E5%8C%96%E5%95%A6%EF%BC%81/6422
**Category:** 資料科學
**Created:** [2023年九月3日 13:46 UTC](https://vip.studycamp.tw/t/%E5%93%88%E4%BD%9B%E6%8A%8A%E5%8D%81%E5%84%84%E7%AF%87%E7%BE%8E%E5%9C%8B%E5%A0%B1%E7%B4%99%E6%96%87%E7%AB%A0%E9%80%9A%E9%80%9A%E9%9B%BB%E5%AD%90%E5%8C%96%E5%95%A6%EF%BC%81/6422 "2023-09-03T13:46:43Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![postman](https://vip.studycamp.tw/user_avatar/vip.studycamp.tw/postman/32/5624_2.png) [@postman](https://vip.studycamp.tw/u/postman)
#### Post date: [2023年九月3日 13:46 UTC](https://vip.studycamp.tw/t/%E5%93%88%E4%BD%9B%E6%8A%8A%E5%8D%81%E5%84%84%E7%AF%87%E7%BE%8E%E5%9C%8B%E5%A0%B1%E7%B4%99%E6%96%87%E7%AB%A0%E9%80%9A%E9%80%9A%E9%9B%BB%E5%AD%90%E5%8C%96%E5%95%A6%EF%BC%81/6422/1 "2023-09-03T13:46:44Z")

</div>

## 事件

作者（ [**鄭紹鈺**](https://www.facebook.com/rainchamber123/posts/pfbid0c5aaZAGVpf9gZz1e8uc5Cc5SCP5jqeTduuDTjb2Hg2FUNzvvUesC1XG5gR9x7Vv9l) ）臉書分享：

我們在哈佛的實驗室，最近釋出了一個全新的「十億級」的文字資料集，原始文本來自1780-1960年美國公有領域的歷史報紙。

透過我們開發的各種深度學習工具，我們提供了超高精準度的電子化文字（已經OCRed），所以這是已經結構並電子化的文字資料集！

重點是------我們 **開源釋出這資料到Hugging Face** 上，以利全世界的人都可以利用！

* * *

## 說明

這計劃的原始影像資料來自美國國會圖書館的Chronicling America檔案庫，我們利用深度學習工具，先是識別了約2,000萬份報紙掃描檔上的11.4億個內容塊，接下來針對標題、文章、作者署名和相關圖片說明。

經由我們特殊的「高效率字母識別模型（EfficientOCR）」來處理成電腦可識別的文字（aka 新細明體或 Times New Roman)。該資料集包含了 4.38 億份已經被賦予結構的報紙文本。

我們也把所有的 Pipeline 開源到 **[GitHub](https://github.com/dell-research-harvard/americanstories)**。我們還創建了開源 Package - LayoutParser 和 EfficientOCR，以幫助研究者可以用上類似的流程來電子化自己有興趣的文本。

這份資料可以協助研究者理解過去美國的歷史變遷，比方說，我們便利用了客製化的Constrasively Trained Contexulized Embeddings，偵測出來了美國歷年來最流行的新聞題目，我們也利用了 Supervised Topic Classifier，從這些新聞資料整理出了許多重要的變數， **可以用作未來經濟研究的迴歸分析** 。

* * *

## 總結

1. 我們釋出了一個十億集的文字資料，是可以用來理解美國過去百年發展最好的資料。

2. 我們釋出了相關的流程跟套件。如果你有想要親自電子化的文檔，也可以從我們的開源工具進一步發展出自己的Pipeline。

3. 一切都是免費跟開源的。

4. 幫老闆感謝一下金主：Harvard Data Science Initiative, Catalyst, and Griffin Fund and MS Azure

* * *

## 資料來源

### [**鄭紹鈺**](https://www.facebook.com/rainchamber123/posts/pfbid0c5aaZAGVpf9gZz1e8uc5Cc5SCP5jqeTduuDTjb2Hg2FUNzvvUesC1XG5gR9x7Vv9l) 臉書

> **[鄭紹鈺](https://www.facebook.com/rainchamber123/posts/pfbid0c5aaZAGVpf9gZz1e8uc5Cc5SCP5jqeTduuDTjb2Hg2FUNzvvUesC1XG5gR9x7Vv9l)**
>
> 〈我們在哈佛把十億篇美國報紙文章通通電子化啦！American Stories---我們lab的『美國新聞故事資料集』釋出來啦！〉...

### Hugging Face

> **[dell-research-harvard/AmericanStories · Datasets at Hugging Face](https://huggingface.co/datasets/dell-research-harvard/AmericanStories)**
>
> We’re on a journey to advance and democratize artificial intelligence through open source and open science.

### GitHub

> **[GitHub - dell-research-harvard/AmericanStories: The official Github for the American Stories...](https://github.com/dell-research-harvard/americanstories)**
>
> The official Github for the American Stories dataset as in {link}

### arXiv

> **[American Stories: A Large-Scale Structured Text Dataset of Historical U.S....](https://arxiv.org/abs/2308.12477)**
>
> Existing full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts from articles, headlines, captions, advertisements, and other layout...

---

<div class="post-metadata">

### Author: ![ChrisWei](https://vip.studycamp.tw/user_avatar/vip.studycamp.tw/chriswei/32/196_2.png) [@ChrisWei](https://vip.studycamp.tw/u/ChrisWei)
#### Post date: [2023年九月5日 13:10 UTC](https://vip.studycamp.tw/t/%E5%93%88%E4%BD%9B%E6%8A%8A%E5%8D%81%E5%84%84%E7%AF%87%E7%BE%8E%E5%9C%8B%E5%A0%B1%E7%B4%99%E6%96%87%E7%AB%A0%E9%80%9A%E9%80%9A%E9%9B%BB%E5%AD%90%E5%8C%96%E5%95%A6%EF%BC%81/6422/3 "2023-09-05T13:10:14Z")

</div>

真棒的 Model~
