Anyone interested in working on mining data?

266 views
Skip to first unread message

Rahul Basu

unread,
Jun 13, 2026, 3:22:22 AMJun 13
to data...@googlegroups.com
Hi friends,

I work on mining in India, primarily Goa. Over time, a lot of data has been collected through various public sources and RTI.  Much of it is in excel, a part is in GIS form. Some of the excel is old government style, needs openrefine to cleanup

Lots of interesting analysis can be done (for example). Some of data includes detailed information on transport of minerals from leases to jetties to ships.

I'm not very good at data manipulation / data visualization and was looking for collaborators. 

If anyone is interested, please ping me.

PS. Odisha seems to have good mining data as well, though much more data collection is needed. Can be looked at in a second phase.

Warmly
Rahul

Dr. C.Gajendran

unread,
Jun 13, 2026, 3:28:04 AMJun 13
to data...@googlegroups.com
Dear Your email find interest to me; please contact me at 9443368980. We shall discuss further.

Regards,

Dr. C. GAJENDRAN  B.E (Civil) M.E (Env.Eng.),Ph.D.,D.R.T.M.,D.I.S.,FIE(Ind.)., IGS., ISG.,
Vice Principal | Director – IQAC | Professor & Head, Civil Engineering | Vice President IIC 
Jawaharlal College of Engineering and Technology, Palakkad

Mobile & Whatsup! : +91-9443368980
skype: c_gajendran
Google Site 
ORCHID : 0000-0002-2765-7075

Twitter : dr_gajendran
Instagram : dr.gajendran


--
Datameet is a community of Data Science enthusiasts in India. Know more about us by visiting http://datameet.org
---
You received this message because you are subscribed to the Google Groups "datameet" group.
To unsubscribe from this group and stop receiving emails from it, send an email to datameet+u...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/datameet/CAN8Wyx_6JXrvcYQnPSw7j%3DKZg2JEWSsB66ur4SRBGaZrsJuDRQ%40mail.gmail.com.

Arun Ganesh

unread,
Jun 13, 2026, 8:58:00 AMJun 13
to data...@googlegroups.com
Hi Rahul, this is a very deep topic and requires many hands to dig through (pun intended). is it possible to share links to some of the collected data so it can be explored? 

Just extracting data from PDF tables requires a lot of effort and it would be best to keep the data open so it can get maximum interest. Or are there reasons to not make this public given its sensitive nature?

bibhu prasad mishra

unread,
Jun 13, 2026, 9:27:42 AMJun 13
to data...@googlegroups.com
Hi Rahul,

I am very interested in collaborating on this. As I am from Odisha, which is a major mineral hub, I believe this is a great opportunity to work with this data.

I would love to help with the data manipulation and visualization, especially as you look toward the second phase involving Odisha. Please let me know how we can get started.


Thanks and Regards
Bibhu Prasad Mishra
Civil Engineer,
B.tech, M.tech.
7894761150, 9937512121

Natasha kalra

unread,
Jun 13, 2026, 11:29:42 AMJun 13
to data...@googlegroups.com, rahul...@gmail.com

Hi Rahul,

Thank you for sharing this. The dataset sounds extremely valuable, especially given the combination of RTI-based information, government records, GIS layers, and transport-chain data linking leases, jetties, and shipments.

I work in the broader sustainability, environmental governance, and data-driven policy space, and I would be interested in understanding the scope of the datasets and the kinds of analyses you are looking to undertake. I am affiliated with the Institute for Social and Economic Change (ISEC), Bengaluru, a premier interdisciplinary research institute engaged in policy-oriented research, consultancy, and capacity-building across areas such as sustainability, environmental governance, urban development, and social transformation. Through our work, we collaborate with government agencies, corporates, development organizations, and communities to develop evidence-based and scalable solutions to contemporary challenges.

It would be great to connect and discuss the available data, current status of data cleaning and organization, and potential avenues for collaboration. Please let me know a convenient time for a call.

Looking forward to learning more.

Warm regards,

Natasha




Dr. Natasha Kalra

Environmental Sustainability Consultant 

Initiator- BaansCraft

https://www.instagram.com/baanscraft/

M: + 91 99167 70495






Rahul Basu

unread,
Jun 14, 2026, 1:54:29 AMJun 14
to datameet
Friends,

Thanks for the interest in this, unexpected for me. I now have to sort through the data and figure out how to organise it for public use. Give me some time on this.

For the moment, here's two datasets I received under RTI around transport permits. 

Background: Goa runs an ore accounting software, which basically accounts for mineral ore at different locations. To move ore from one point to another, a transport permit is required. This is usually for a decent quantity, say 10,000 tons, to move from Mine X to Jetty Y. Then there are truck trips (usually 10 tons each), and the trucks are monitored real time. There are also trips by barge from the jetty to the ship, and some rail as well. When all the trips against the permit are completed, there's usually some difference, which gets adjusted. The software provider has changed from Megasoft to GEL. GEL software is called Bhumija.

Mining was in three major phases: pre-2012 ban; 2015-2018; and 2024 onwards. Ore is mined in Goa, most exported, some consumed in Goa, some shipped by barge to JSW Dolvi, Maharashtra (recent). Ore is also imported into Goa from overseas. Some ore mined in Karnataka and Maharashtra also moves through Goa, mostly for export but occasionally for use in steel plants in Goa. 

The first folder has transport permits. Lots of them. From 1-Apr-2020-2025

The second folder is around closing stock also has some version of the manuals for the software. I had asked for:

a) Manuals of the Bhumija software
b) Guidelines / SOPs / instruction for monitoring ore through Bhumija 
c) In excel format, for each major mineral ore transport trip after 25-Mar-2024, permit number, carrier registration number, the source location with time and weight at the source weighbridge and the destination location with time and weight 
d) In excel format, closing stock of major mineral ore as of 31-Mar-2024 and 31-Mar-2025 for all locations in Bhumija

I have a few questions from the data:
1) Does the ore balance at any location go negative at any point in time?
2) Overall statistics - major input locations, major transport routes, total quantities, etc.
3) Is the data complete? Is any further information needed / desireable?

More questions will come as other data sets get added

Please feel free to work with this while I try to set up a better data repository. Email me directly if you have questions on the data in the two folders.

Warmly
Rahul

Sanjay Bhangar

unread,
Jun 14, 2026, 4:19:46 AMJun 14
to data...@googlegroups.com
Hi Rahul,

Thank you so much for this most interesting dataset - was a fun thing to wake up to Sunday morning and couldn't resist spending a bit of time with AI tools to clean up the Excels a bit, ingest it into a database, and enable some basic querying and visualization. I've been wanting to play around with Datasette for some time now, and this seemed like a good use-case to get up something simply to explore the data.

 - I have put up what I came up with at https://goamines.whydidweevendothis.com/
 - List of tables, query executor and some example queries: https://goamines.whydidweevendothis.com/goamines
 - A map showing route arcs of movements: https://goamines.whydidweevendothis.com/static/routes_map.html (geolocation of points is currently imprecise)

All the code for the data ingestion, including notes made by my AI agent along the way can be found at https://github.com/batpad/goamines

WARNING: I have really not spent a lot of time verifying ingestion scripts, etc. so there could be discrepancies in the data, especially since there was quite a lot of inconsistency in the source data - please do not treat this as an accurate data source without doing your own verification.

I likely will not have a lot of time to continue working on this, but it is a really interesting data-set and it was really fun to spend some time building something with it. If folks find it useful / want to take it over / move the repo or domain somewhere etc. please just let me know and of course feel free to make Issues and Pull Requests on the repository. Also, if anyone wants the SQLite file (around 90MB) just let me know and I'll put it somewhere for download.

Thanks again Rahul - this was some really amazing work gathering this data - full credit to you, and I hope you find this experiment in processing and visualizing it useful.

Cheers,
Sanjay

Ashim Jain

unread,
Jun 15, 2026, 8:33:06 AMJun 15
to data...@googlegroups.com
Friends, 

Just catching up on this thread.

Out of the huge datasets, if specific evidence can be extracted showing govt. and/or corporate culpability in mineral exploitation, how can we make this easy to understand and reach the affected masses in those areas?  (Because it's only the knowledge and power of the masses that has the potential to change the status quo.)

For decades, we have been reading about the tribal and rural populations' struggle against being dispossessed of the land and resources they have nurtured for thousands of years.  It is anti-democratic and further leads to concentration of wealth and power, exacerbating the wealth disparities that are causing havoc.

In most cases, it is being done illegally (manipulating the gram sabhas, under the table deals, etc.).  

I'm willing to do my part (web app development and contribution to a combined effort) if we can form a team.

- Ashim, Udaipur, Rajasthan


Rahul Basu

unread,
Jun 15, 2026, 10:13:54 AMJun 15
to data...@googlegroups.com
The data exists, every week there are new reports of illegal mining or illegalities in mining. 

Communication to mobilize masses is hard, but possible. Consider how there were crowds on the streets during IAC in 2012 (about corruption in coal blocks among others), which ended up with the NDA government in Goa in 2012 and nationally in 2014. 

We have a different approach which we tried through the Goenchi Mati Movement back in 2017. See the manifesto we wanted politicians to sign up to. It was very hard, we didn't succeed in changing the politics. Here's a 13 minute video of what happened. 

This is a summary of what we think would be fair and can be sustained. The key advocacy point for us has been the government as our trustee, not so much vilifying the miners as the problem is really a systemic one - everyone in the present generation is trying to capture some part of the shared inheritance of our children & future generations. Note that the target of our communication is the unaffected masses, neither benefiting from mining nor facing the ground consequences - the pitch is that their children's inheritance is being stolen and their government corrupted from the proceeds; but there's an alternative that is better without completely banning mining.

At the moment, if we find irregularities, we approach the government, and failing that, the courts. Even the Supreme Court is finding it hard to control illegal sand mining in the Chambal river. So yes, we do need to communicate better.

Warmly
Rahul
Today is the first day of the rest of your life !

You received this message because you are subscribed to a topic in the Google Groups "datameet" group.
To unsubscribe from this topic, visit https://groups.google.com/d/topic/datameet/WLHGLnl5IX0/unsubscribe.
To unsubscribe from this group and all its topics, send an email to datameet+u...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/datameet/CAPaFkDyWah8H8r_bw-X%3DU5myyZFYyKEVD9ohGZi3posYipD_gQ%40mail.gmail.com.

Bryan

unread,
Sep 1, 2026, 11:20:55 PM (8 days ago) Sep 1
to datameet
Hi Rahul and Sanjay,

I worked through the reconstruction and prepared a concrete correction, with the before/after file and reproducible tests:
https://github.com/batpad/goamines/issues/1

Two date boundaries affect the result: movements on 31 March 2024 were being added to that day's closing-stock baseline, while timestamped movements on 31 March 2025 were excluded by the upper-bound comparison. The patch starts after the opening closing-stock day and includes the whole final day.

On the downloaded public database, four reconstructed balances change. All 166 patched reconstructions match an independent calendar-date calculation. The issue contains the four-row report and a free download with the patch, ten boundary tests, full receipt and instructions. The test runs in memory and leaves the source database unchanged.

This uses an end-of-day closing-stock interpretation. It does not resolve missing baselines or source completeness, and negative values alone do not establish physical shortage or wrongdoing. The royalty stream still has no 2024 baseline.

Thank you for making the data and code available — happy to discuss the cutoff convention or adapt the patch.
Bryan / CivicDataForge

Sanjay Bhangar

unread,
Sep 2, 2026, 1:43:25 AM (8 days ago) Sep 2
to data...@googlegroups.com
Hey Bryan,

Thanks for this - will test and update the code and site soon.

Would love to know more about your project and methodology. It's clear from your messages and your Github account that this is some Human / AI collaboration experiment. To be clear, I appreciate your messages and fixes, but in the spirit of community and transparency, would really appreciate a quick introduction and what you're trying to do here - are you trawling multiple open data lists and having AI scan for issues and help out? Is it just datameet? Is this part of some organized effort? Some background would really help - Datameet is also a community of human beings, and it's helpful to have some background and some context of the human being(s) behind AI-assisted work.

If you plan on contributing more regularly to the list, it would be really great if you could start a new thread with a brief introduction of your work and methodology and how you're using AI to make contributions, etc. Ofc, this is not a requirement, and your work is appreciated - I just feel in this world of increasingly AI-generated text, it's really nice to have a sense of the human intention behind some of this - otherwise we will end up with mailing lists with agents speaking to agents (which maybe ok?).

Thanks again for your message and contribution - if you are comfortable sharing, would really appreciate a bit of background around the work you're doing.

Cheers,
Sanjay

--
Datameet is a community of Data Science enthusiasts in India. Know more about us by visiting http://datameet.org
---
You received this message because you are subscribed to the Google Groups "datameet" group.
To unsubscribe from this group and stop receiving emails from it, send an email to datameet+u...@googlegroups.com.

Bryan

unread,
Sep 2, 2026, 3:40:56 AM (8 days ago) Sep 2
to datameet
Hi Sanjay — thank you for asking directly. I’ve posted a separate introduction explaining who I am, what CivicDataForge does, and how our human–AI workflow operates. I appreciate the welcome and the reminder that human context matters. — Bryan

Rahul Basu

unread,
Sep 2, 2026, 5:26:52 AM (8 days ago) Sep 2
to data...@googlegroups.com
Friends, esp Sanjay & Bryan,

I've organised the data better. There's a bunch of sub-folders in this folder. In each folder, I've tried to create a doc file that provides some background.

The mining transport data that Sanjay worked on is in Mining Transport - Pub. I've added some older data to this folder, as well as newer data (could be some overlap in periods).

BTW, the work around the AQI data logger by Arun is basically air quality data from stations established specifically to monitor air quality along these transportation routes.

Please point out any additional information that may be useful for analysis.

Looking forward to the insights of the community.

Rahul
Today is the first day of the rest of your life !

On Sun, Jun 14, 2026 at 11:24 AM Rahul Basu <rahul...@gmail.com> wrote:
You received this message because you are subscribed to a topic in the Google Groups "datameet" group.
To unsubscribe from this topic, visit https://groups.google.com/d/topic/datameet/WLHGLnl5IX0/unsubscribe.
To unsubscribe from this group and all its topics, send an email to datameet+u...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/datameet/c67f7584-fcbc-4a1c-83ba-dff8b0e08085n%40googlegroups.com.

Rahul Basu

unread,
Sep 4, 2026, 3:03:12 AM (6 days ago) Sep 4
to data...@googlegroups.com
On mining transport, there's a big set of data as part of the Legislative Assembly Questions. Basically all the trips for major mining leases, routes, etc. It is in multiple PDFs, would greatly appreciate it if anyone can process these into databases.
http://static.goavidhansabha.gov.in/goalpub/vr/167_01092026.zip

Thanks in advance.

Rahul
Today is the first day of the rest of your life !

Bryan

unread,
Sep 4, 2026, 6:13:11 PM (5 days ago) Sep 4
to datameet
Hi Rahul and everyone,

I've converted the LAQ 167 archive into a SQLite database, with compressed JSON records and smaller JSON tables alongside it:

https://drive.google.com/file/d/1Npw5WK7eOCr7hKzRiirjT0lgBGLDRcKS/view

The ZIP is about 31 MB. Download and extract it, then open goa-mining.sqlite in a SQLite viewer and select the trip_evidence view. The README explains the tables, and QUERIES.sql includes starting queries. This is a free community contribution, with no paid service needed to use the file.

What's in it:
- 363,507 individual trips from all ten Annexure D PDFs, covering 55,420 pages.
- All 798 published daily route totals, plus block/lessee, monthly production, dispatch and route-permission context from A/B/C.
- Source PDF names, page references and hashes, so a record can be traced back to the original document.

The trip-level counts and ore quantities reconcile exactly with all 798 published daily route totals. There are no duplicate trip IDs or parsing exceptions. I also checked every cell in 290 rows across 36 sampled pages with a second PDF parser; that independent check is sampled, not every page.

The actual trip dates range from 6 July 2024 to 19 August 2026, with different cutoffs by block. The September answer date is not a claim that every block is current to September. Block II has contextual records but no trip PDF in Annexure D.

All trips match a published route-permission entry using the route text, without fuzzy matching. The package also flags a 0.01 difference in one stated production total and four differing commencement-date statements across annexures, retaining both versions. The printed clocks lack AM/PM/timezone, so I haven't turned them into precise timestamps or claimed hourly violations. AQI remains a linked companion source, not an inferred route-to-sensor or causal match.

If a row looks wrong, a PDF/page reference or trip ID will let us check it precisely. I hope this saves the group a substantial amount of preparation work.

Bryan
CivicDataForge
civicdat...@gmail.com

Arun Ganesh

unread,
Sep 4, 2026, 10:41:31 PM (5 days ago) Sep 4
to data...@googlegroups.com

Rahul Basu

unread,
Sep 4, 2026, 10:58:58 PM (5 days ago) Sep 4
to data...@googlegroups.com
Wow. That's amazing. And so quick.Thanks a ton Bryan.

Rahul
Today is the first day of the rest of your life !

On Sat, Sep 5, 2026 at 3:43 AM Bryan <neoa...@gmail.com> wrote:

Rahul Basu

unread,
Sep 5, 2026, 1:20:17 AM (5 days ago) Sep 5
to data...@googlegroups.com
Hi Bryan

All the time stamps are Indian Standard Time.

Rahul
Today is the first day of the rest of your life !

Bryan

unread,
Sep 5, 2026, 3:27:12 AM (5 days ago) Sep 5
to data...@googlegroups.com
Hi Rahul and everyone,

Thanks for clarifying the timezone, Rahul, and thank you both for the kind words.

I've added another useful layer to the database: AM/PM and full local timestamps for 330,118 of the 363,507 trips, using the related exports already in the shared collection. Indian Standard Time is now explicit, with UTC equivalents included for anyone combining this with other time-series data.

The enriched edition is here:
https://drive.google.com/file/d/1EpSQ4hrUt3Nw80nETayj4ip9N4f7dMfS/view

This is the full standalone package, about 42 MB. Download and extract it, open goa-mining.sqlite, and use the same trip_evidence view as before. The added fields are also in the JSON export, and the README explains them. No separate merge is needed.

All original trips, quantities, source references and supporting tables are retained. The remaining 33,389 trips still have their original dates and printed clocks; unresolved AM/PM is left blank rather than guessed. Daily totals remain unchanged. The earlier download remains available too.

I hope this makes the next stage of analysis easier. Happy to help with useful next questions as the group explores it.

Bryan
CivicDataForge
civicdat...@gmail.com

Reply all
Reply to author
Forward
0 new messages