Dear Dr. Anthony,Thank you for taking the time to answer our questions and for developing AntConc software. I am using AntConc 4.3.1 for my PhD thesis which is about lexical bundles.
I loaded my corpus into AntConc (TXT files after cleaning). I used the following criteria (N gram size: 4, Min. frequency: 7, Min Range: 19). The extraction generally produced the expected results and I applied the overlap subsumption rules for some bundles (Chen & Baker, 2010).
However, I got some results that I am not sure why. For example, I got the following two bundles: of
the operation the
and commander’s intent the. I excluded them from my final bundle list because I considered them artifacts of automatic extraction and I explained this decision in my data analysis section of my thesis.
(of
the operation the)
appeared in the KWIC lines as of the operation. The…, and (commander’s
intent the) occurred as commander’s intent. The…
My supervisor commented that this was unsual because corpus tools are generally set up to locate patterns within a specific sentence boundary (e.g. full stop, question marks, exclamation marks ... etc).
Could you please clarify whether the N-Gram tool in AntConc can generate n-grams across sentence boundaries and whether there is a setting that can prevent this? I need to determine whether these sequences reflect normal extraction behaviour or a problem with my settings so that I can revise my Results chapter accurately.
Best regards,
Amal