Sentence Boundary Cases - AntConc Results

30 views
Skip to first unread message

Amal Alzahrani

unread,
Sep 16, 2026, 3:04:48 AM (6 days ago) Sep 16
to AntConc-Discussion
Dear Dr. Anthony,
Thank you for taking the time to answer our questions and for developing AntConc software. I am using AntConc 4.3.1 for my PhD thesis which is about lexical bundles. 

I loaded my corpus into AntConc (TXT files after cleaning). I used the following criteria (N gram size: 4, Min. frequency: 7, Min Range: 19). The extraction generally produced the expected results and I applied the overlap subsumption rules for some bundles (Chen & Baker, 2010).

However, I got some results that I am not sure why. For example, I got the following two bundles: of the operation the and commander’s intent the. I excluded them from my final bundle list because I considered them artifacts of automatic extraction and I explained this decision in my data analysis section of my thesis.
(of the operation the) appeared in the KWIC lines as of the operation. The…, and (commander’s intent the) occurred as commander’s intent. The…
My supervisor commented that this was unsual because corpus tools are generally set up to locate patterns within a specific sentence boundary (e.g. full stop, question marks, exclamation marks ... etc). 

Could you please clarify whether the N-Gram tool in AntConc can generate n-grams across sentence boundaries and whether there is a setting that can prevent this? I need to determine whether these sequences reflect normal extraction behaviour or a problem with my settings so that I can revise my Results chapter accurately.

Best regards,
Amal

Laurence Anthony

unread,
Sep 16, 2026, 3:10:01 AM (6 days ago) Sep 16
to ant...@googlegroups.com
Hi Amal,

As AntConc doesn't perform an automatic sentence boundary detection on your corpus, it cannot search for N-grams within sentence boundaries by default. However, my guess is that you just need to select the following option in the N-Gram tool.

image.png

I hope that helps!

Laurence.


###############################################################
Laurence ANTHONY, Ph.D.
Professor of Applied Linguistics
Faculty of Science and Engineering
Waseda University
3-4-1 Okubo, Shinjuku-ku, Tokyo 169-8555, Japan
E-mail: antho...@gmail.com
WWW: http://www.laurenceanthony.net/
###############################################################


--
You received this message because you are subscribed to the Google Groups "AntConc-Discussion" group.
To unsubscribe from this group and stop receiving emails from it, send an email to antconc+u...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/antconc/9a97b0bf-4bda-4832-b02f-e2a805a49697n%40googlegroups.com.

Amal Alzahrani

unread,
Sep 16, 2026, 3:58:44 AM (6 days ago) Sep 16
to AntConc-Discussion
I used the option in the screenshot and my results changed. I got 190 lexical bundles instead of 238. Does this mean my previous results were wrong or should I just exclude the two bundles that were boundary cases bacuse I already analyzed the previous results?

Laurence Anthony

unread,
Sep 16, 2026, 4:02:13 AM (6 days ago) Sep 16
to ant...@googlegroups.com
Hi,

>I used the option in the screenshot and my results changed. I got 190 lexical bundles instead of 238. Does this mean my previous results were wrong or should I just exclude the two bundles that were boundary cases bacuse I already analyzed the previous results?

The previous results were not wrong, but they just mean different things to new results. The new results are for N-grams separated only by spaces. The old results included any kind of punctuation between the words.

The important thing is to decide which type of N-gram you want to use.

Regards,

Laurence.

###############################################################
Laurence ANTHONY, Ph.D.
Professor of Applied Linguistics
Faculty of Science and Engineering
Waseda University
3-4-1 Okubo, Shinjuku-ku, Tokyo 169-8555, Japan
E-mail: antho...@gmail.com
WWW: http://www.laurenceanthony.net/
###############################################################

Amal Alzahrani

unread,
Sep 16, 2026, 8:18:11 AM (6 days ago) Sep 16
to AntConc-Discussion
Thank you again. You're really helping me alot. 

When I applied the new settings, I lost some useful and meaningful bundles. I am thinking to keep my original AntConc settings and mention in the methodology that the option "Only include ngrams with whitespace separators" was tested but it was not adopted because it also excluded within-sentence sequences containing punctuation and that the two bundles that crossed sentence boundaries were identified through manual checking of the KWIC concordance lines, and they were excluded from the final bundle list. I actually checked every single bundle in their contexts (concordance lines) because I did structural and functional classifications. 

Laurence Anthony

unread,
Sep 16, 2026, 8:22:25 AM (6 days ago) Sep 16
to ant...@googlegroups.com
Hi Amal,

Yes, that sounds like a very good strategy. I can imagine many cases where an interesting N-gram might cross a sentence boundary. For example:
That's true. But, what about ...
Stop. Danger.

Good luck with your research!

Laurence.

###############################################################
Laurence ANTHONY, Ph.D.
Professor of Applied Linguistics
Faculty of Science and Engineering
Waseda University
3-4-1 Okubo, Shinjuku-ku, Tokyo 169-8555, Japan
E-mail: antho...@gmail.com
WWW: http://www.laurenceanthony.net/
###############################################################

Reply all
Reply to author
Forward
0 new messages