Publishing Partner: Cambridge University Press CUP Extra Publisher Login
amazon logo
More Info

New from Oxford University Press!


Oxford Handbook of Corpus Phonology

Edited by Jacques Durand, Ulrike Gut, and Gjert Kristoffersen

Offers the first detailed examination of corpus phonology and serves as a practical guide for researchers interested in compiling or using phonological corpora

New from Cambridge University Press!


The Languages of the Jews: A Sociolinguistic History

By Bernard Spolsky

A vivid commentary on Jewish survival and Jewish speech communities that will be enjoyed by the general reader, and is essential reading for students and researchers interested in the study of Middle Eastern languages, Jewish studies, and sociolinguistics.

New from Brill!


Indo-European Linguistics

New Open Access journal on Indo-European Linguistics is now available!

Query Details

Query Subject:   WebCorp Concordance Counts
Author:   Jerry Kurjian
Submitter Email:  click here to access email

Linguistic LingField(s):  Computational Linguistics
Text/Corpus Linguistics

Query:   Hi all,
I have a question about the concordance counts produced by the WebCorp site:

For example, if I search ''suggest you don't'' vs. ''suggest that you
don't'' using WebCorp (via Google) I get, at the bottom of the page, a
concordance count of 187 vs. 96 kwics respectively. However, if I search
the same two terms, in quotes, on Google, I get 34,200 vs. 16,200 hits.
The ratios are similar though not the same.

Does anyone have insight into how WebCorp calculates/filters its
concordances or why these two engines are so different in the number of
hits they return?

In fact, it is nice to have the more manageable number produced by WebCorp,
and the external collocate counts it creates. But if I am interested in
the frequency of ''I'' collocating with the two search terms based on
WebCorp, I'd like to be clearer how those two counts are derived.

LL Issue: 16.1291
Date posted: 22-Apr-2005


Sums main page