Google Groups no longer supports new Usenet posts or subscriptions. Historical content remains viewable.
Dismiss

SCDB: db put failed (solaris 8?)

0 views
Skip to first unread message

David Powicki

unread,
Jun 19, 2001, 12:21:34 PM6/19/01
to

Has anyone reported similar problems with Solaris 8 and v6.0?

I've looked in google's archives, but would like to know the current
state of affairs with Solaris 8 and V6.0. Right now it still looks too
risky to move away from 5.2.

Thanks

David


>
> PMDF version is PMDF V6.0-23
> libpmdf.so version V6.0-025; linked 17:32:38, Jun 7 2001
> SunOS mail-gw2 5.6 Generic_105181-23 sun4u sparc SUNW,UltraSPARC-IIi-cEngine
>
> After applying the June 8 libpmdf.so, our problems got worse. We are
> moving to Iplanet's Messaging Server 5.1 in light of these issues (IMS
> containing the PMDF code we've grown to love as well as similar config
> files).
>
> We are told that 6.1 might not even fix these problems (this is by
> Process support, which we pay a good amount of money to). In the
> interim, we will be downgrading from 6.0 to version 5 (which we've used
> previously, with no problems).
>
> Has anyone experienced any problems with downgrading? I'm guessing the
> best way would be to back up pmdf.cnf and mappings (and some others) and
> just uninstall PMDF 6.0 completely. Then begin anew with PMDF 5.
> Nobody has ever had any problems with PMDF 5's database being corrupted,
> right?
>
> PMDF 5 was a great product and Innosoft's support was the best support
> I've ever received (on anything, those people were great). It's
> unfortunate that we need to change platforms from this inability to
> solve some serious problems.
>
> Thanks,
> Jeff
>
> Jeff Beck
> Academic Computing // 212.241.7091
> Mount Sinai School of Medicine
>
> "Danger lies not in what we don't know, but in what
> we think we know that just ain't so." - Mark Twain
>
> ----- Original Message -----
> From: David Richards <d.ric...@qut.edu.au>
> Date: Thursday, June 14, 2001 5:03 am
> Subject: Re: SCDB: db put failed
>
> > > If this problem could be fixed, I would be eternally gratefull.
> > So far
> > > this week my pager has gone off twice in the middle of the night
> > related> to the queue databases being corrupted. I've been using
> > PMDF for quite
> > > a few years, and generally I think it's a great product, but
> > this type of
> > > problem has definitely made me entertain thoughts of looking at
> > alternative> products.
> >
> > Hey, you too? My mobile is continously going off middle of the day,
> > middle of the night, bloody all the time!
> >
> > I have gotten to the stage where I have the laptop set up next to my
> > bed, I can reach over, run a script and get it running again.
> >
> > I have registered my interest in being a beta site for 6.1, but not
> > heard anything yet ... could not be much worse than 6.0!
> >
> > Dave.
> >
> > >
> > > John Meyers
> > > Computing Services
> > > Wright State University
> > > Dayton, Ohio.
> > > E-Mail: john....@wright.edu
> >

--

David Powicki Network Analyst/Postmaster OIT Network Services
Voice: 413.545.1605 Fax: 413.545.3203 University of Massachusetts
email: dpow...@nic.umass.edu Amherst, MA 01003-4640

Steve J Depinet

unread,
Jun 19, 2001, 12:34:18 PM6/19/01
to

# pmdf ver

PMDF version is PMDF V6.0-23
libpmdf.so version V6.0-24; linked 15:02:02, Mar 8 2001
SunOS mailgate1 5.8 Generic_108528-04 sun4u sparc SUNW,Ultra-80

Definitely happens with Sol 8 and PMDF 6.0.

High hopes for PMDF v6.1 (but too busy building another system to
serve as a Beta tester).

I considered installing v5.2 on the second server, but the support and
other issues overrode the stability of 5.2.

Steve de Pinet Northern Arizona University
Systems Programmer, Sr. Box 5100
Information Technology Services Flagstaff, Az 86011
Steve....@nau.edu (520) 523-6843
http://jan.ucc.nau.edu/~sjd

Hunter Goatley

unread,
Jun 19, 2001, 12:33:56 PM6/19/01
to
> Has anyone reported similar problems with Solaris 8 and v6.0?

The data that I've received would indicate no, but that's not
definitive. This question has been asked repeatedly, but no one ever
replies. I assume that means few, if any, PMDF customers are running
Solaris 8.

> I've looked in google's archives,

Just a reminder that the Info-PMDF archives are available for
searching from:

http://www.pmdf.process.com/

and specifically

http://www.pmdf.process.com/scripts/mxarchive/as_init.com?Info-PMDF

> but would like to know the current
> state of affairs with Solaris 8 and V6.0. Right now it still looks too
> risky to move away from 5.2.

On April 23, I posted a summary of my understanding of the problems
and versions, etc. You can find it here:

http://www.pmdf.process.com/scripts/mxarchive/archive_search.com?TEXT=R419364-422100-mail%24archives%3A%5Binfo-pmdf%5Dinfo-pmdf.2001-04

The most I know about Solaris 8 comes from someone at Innosoft:

> - Another Innosoft person states that Sun/iPlanet doesn't support
> Solaris 7, and he's not aware of SleepyCat issues on Solaris 8


Hunter
------
Hunter Goatley, Process Software, http://www.process.com/
<goath...@GOATLEY.COM> http://www.goatley.com/hunter/

Hunter Goatley

unread,
Jun 19, 2001, 12:38:44 PM6/19/01
to
> Definitely happens with Sol 8 and PMDF 6.0.

I stand corrected. I'm not very surprised; the problems seems to be
in the SleepyCat code itself (and/or how PMDF is using that
code---these are the areas we're investigating).

Steve J Depinet

unread,
Jun 19, 2001, 12:46:03 PM6/19/01
to

On Tue, 19 Jun 2001, Hunter Goatley wrote:

> > Has anyone reported similar problems with Solaris 8 and v6.0?
>
> The data that I've received would indicate no, but that's not
> definitive. This question has been asked repeatedly, but no one ever
> replies. I assume that means few, if any, PMDF customers are running
> Solaris 8.

Hunter, I just brought this system online, and have experienced DB
corruption problems. I hadn't noticed that none of the previous posts on
this issue had come from Sol 8 systems, in fact, I was under the
impression that there were a few reports from admins if Sol 8 systems.

If there's any debug output that you'd like to see, let me know what
options to set, I'd be happy to forward it to you.

>
> > I've looked in google's archives,
>
> Just a reminder that the Info-PMDF archives are available for
> searching from:
>
> http://www.pmdf.process.com/
>
> and specifically
>
> http://www.pmdf.process.com/scripts/mxarchive/as_init.com?Info-PMDF
>
> > but would like to know the current
> > state of affairs with Solaris 8 and V6.0. Right now it still looks too
> > risky to move away from 5.2.
>
> On April 23, I posted a summary of my understanding of the problems
> and versions, etc. You can find it here:
>
> http://www.pmdf.process.com/scripts/mxarchive/archive_search.com?TEXT=R419364-422100-mail%24archives%3A%5Binfo-pmdf%5Dinfo-pmdf.2001-04
>
> The most I know about Solaris 8 comes from someone at Innosoft:
>
> > - Another Innosoft person states that Sun/iPlanet doesn't support
> > Solaris 7, and he's not aware of SleepyCat issues on Solaris 8
>
>

Steve J Depinet

unread,
Jun 19, 2001, 12:54:13 PM6/19/01
to
In our case, the problem manifests itself in the queue_cache DB becoming
corrupted, and the queues filling up with messages that aren't being
tried for delivery. If we shut down PMDF, do a 'pmdf cache -rebuild' and
restart PMDF, the queues start to deliver email again. I often manually
run the channels affected, as well, to catch up with the backlog sooner.

On one occasion, another admin had begun this process, and I (not knowing
that he was doing so) started the cache -rebuild. We started getting the
SCDB: dbput failed messages, and I ended up rebooting the machine (with
pmdf disabled), renaming every entry in the queue_cache directory, and
rebuilding the queue_cache DB. Then restarting PMDF. This resolution
worked for that case. But most of the time, we just use the stop, rebuild,
restart resolution.


Steve de Pinet Northern Arizona University
Systems Programmer, Sr. Box 5100
Information Technology Services Flagstaff, Az 86011
Steve....@nau.edu (520) 523-6843
http://jan.ucc.nau.edu/~sjd

On Tue, 19 Jun 2001, Hunter Goatley wrote:

> > Definitely happens with Sol 8 and PMDF 6.0.
>
> I stand corrected. I'm not very surprised; the problems seems to be
> in the SleepyCat code itself (and/or how PMDF is using that
> code---these are the areas we're investigating).
>

Hunter Goatley

unread,
Jun 19, 2001, 1:36:33 PM6/19/01
to
> In our case, the problem manifests itself in the queue_cache DB becoming
> corrupted,

Is there anything you're doing management-wise that precedes such
instances? The SleepyCat stuff is very susceptible to problems when
its temporary files are messed with, deleted, etc. I'm just trying to
make sure there's not some cron job doing something that could
interfere with what PMDF is doing.

We're doing testing on PMDF itself to make sure that PMDF itself is
using SleepyCat in a manner that will keep SleepyCat happy.

> But most of the time, we just use the stop, rebuild,
> restart resolution.

Which, from what I've read, seems to always solve the problem, for the
time-being.

Thanks!

Richard Loken

unread,
Jun 19, 2001, 2:19:26 PM6/19/01
to
It would be interesting to know if Esys (Messaging Direct?) and Cyrus Imap
are having database problems with SleepyCat.

I also speculate that these problems are more likely to show up with increased
volumes of traffic. After I fixed my kernel quota parameters I have not seen
any SCDB errors and the worst time for such errors was when the system was
hung over the Easter weekend and we had a backlog of thousands of messages,
I spent the morning rebuilding and recorrupting the cache until the backlog
was gone and then the system hehaved quite well again.

So... With 10,000 messages a day the problems are minimal or nonexistant
when running Tru64 Unix on a DS20E but that may not the case running
50,000 messages a day on Alpha 2100 with 512Myte of memory. ???
---
Richard Loken VE6BSV, Systems Programmer - VMS
Athabasca University
Athabasca, Alberta Canada
** rich...@admin.athabascau.ca **

Steve J Depinet

unread,
Jun 19, 2001, 2:50:55 PM6/19/01
to

On Tue, 19 Jun 2001, Hunter Goatley wrote:

> > In our case, the problem manifests itself in the queue_cache DB becoming
> > corrupted,
>
> Is there anything you're doing management-wise that precedes such
> instances? The SleepyCat stuff is very susceptible to problems when
> its temporary files are messed with, deleted, etc. I'm just trying to
> make sure there's not some cron job doing something that could
> interfere with what PMDF is doing.

Nothing I can think of. When the corruption occurs, it's not noticed until
the (locally written) mail-check system sends a page (which isn't
dependant on the pager channel). I'll keep my eyes open, though.

>
> We're doing testing on PMDF itself to make sure that PMDF itself is
> using SleepyCat in a manner that will keep SleepyCat happy.
>
> > But most of the time, we just use the stop, rebuild,
> > restart resolution.
>
> Which, from what I've read, seems to always solve the problem, for the
> time-being.
>
> Thanks!
>
> Hunter
> ------
> Hunter Goatley, Process Software, http://www.process.com/
> <goath...@GOATLEY.COM> http://www.goatley.com/hunter/
>

Steve de Pinet Northern Arizona University

Larry M. Rosenbaum

unread,
Jun 19, 2001, 3:06:46 PM6/19/01
to
>On Tue, 19 Jun 2001, Hunter Goatley wrote:
>
>> > In our case, the problem manifests itself in the queue_cache DB becoming
>> > corrupted,
>>
>> Is there anything you're doing management-wise that precedes such
>> instances? The SleepyCat stuff is very susceptible to problems when
>> its temporary files are messed with, deleted, etc. I'm just trying to
>> make sure there's not some cron job doing something that could
>> interfere with what PMDF is doing.
>
>Nothing I can think of. When the corruption occurs, it's not noticed until
>the (locally written) mail-check system sends a page (which isn't
>dependant on the pager channel). I'll keep my eyes open, though.

In our case, it would happen only after a program used "sendmail -bs" to send mail (and not every time).

If the hang involved "SCDB: db put" errors, PMDF had to be removed from memory to fix it:

ipcs -a | grep pmdf | awk '{printf "-m %d\n",$2}' | xargs ipcrm


--
========================
Larry M. Rosenbaum rosen...@ornl.gov
Bldg 4500-N, Room E-218 865 574-8155 phone
PO Box 2008, MS 6271 865 241-4000 fax
Oak Ridge, TN 37831-6271

Oak Ridge National Laboratory, Network Computing Services group

David Richards

unread,
Jun 19, 2001, 7:09:40 PM6/19/01
to
I realise you guys want to go with a 'standard' database system, and
SleepyCat performs very well .... normally. But, wouldn't it be better
for your reputation and product stability to go either back to the
databases used in 5.2 or another standard like ndbm or gdbm or
something??

I would really love a stable PMDF, I really would.

Dave.

--
David Richards
Project Manager (Messaging)
Information Technology Services
Queensland University of Technology

Hunter Goatley

unread,
Jun 19, 2001, 7:10:32 PM6/19/01
to
> I realise you guys want to go with a 'standard' database system, and
> SleepyCat performs very well .... normally. But, wouldn't it be better
> for your reputation and product stability to go either back to the
> databases used in 5.2 or another standard like ndbm or gdbm or
> something??

> I would really love a stable PMDF, I really would.

So would we, Dave. I should probably avoid rambling replies like this
one in a public forum, but it was asked publicly, so....

First, as I've discussed before, we aren't licensed to do anything
with pre-V6.0 code. Having said that, we in Engineering have actually
discussed trying to go back to what V5.2 used. Licensing issues
aside, the PMDF V6.0 code base is radically different from V5.2, and
trying to go back would undoubtedly take more work than stabilizing
the SleepyCat code. (It might be worth reminding everyone that going
from V5.2 to V6.0 entailed massive changes to PMDF, including adding
SleepyCat, porting most of the Pascal code to C, adding MessageStore,
porting PMDF to NT, and more. The code changed a lot, and I don't
even know if the old code *could* be used in the current code.)

From what they've told me, Innosoft found that the code they were
using did not scale well for sites with higher throughput. So they
decided to replace the home-grown code in favor of something better.
They chose SleepyCat. Given that I wasn't involved with PMDF at the
time, I don't know exactly why they did choose SleepyCat, but I'm sure
they wouldn't have gone that route if they'd believed ndbm or gdbm or
something else had been better. (In fact, from the web searches we've
done on the topic, the consensus in the UNIX world seems to be that
SleepyCat is the best, even though it's far from perfect.)

We've looked at the code with the thought of going back to the
home-grown stuff, or replacing it with something else, but it would
involve substantial rewriting (again) of the code, an effort that
would take months, at least. So we're trying our best to make the
best of the situation and stabilize PMDF's use of SleepyCat. As I've
said, V6.1 will have the latest SleepyCat code. We're also working on
improving how PMDF uses and interfaces with the SleepyCat code to
better stabilize the product. We have some suggestions from the
Innosoft folks for changes that might help. The SleepyCat stuff seems
to work well when there's only one thing at a time going on. But PMDF
is often doing several things at a time, and when processing lots of
messages, that seems to be where things go south. That's what we're
trying to address, now that we've moved to the latest SleepyCat
source.

So, rather than force the UNIX customers to suffer through the
problems for another year or so while we drop the SleepyCat stuff and
try something else, we're trying to concentrate on stabilizing what's
there as quickly as we can. It's a frustrating situation for all of
us (the problems existed, and went unaddressed, long before PMDF was
transferred to Process Software), but we're working quite diligently
on stabilizing the inherited problems as quickly as we can.

Scott Wood

unread,
Jun 20, 2001, 10:27:11 AM6/20/01
to
My experiences also lead me to believe that the SleepyCat problems are pronounced
under higher volumes of traffic. We've actually been running pmdf 6.0 under
Solaris 8 (on an E450, two cpus, one gb memory) since March and things are
generally pretty stable especially now that the summer is here and our mail
traffic is reduced. During the school year our volume of mail messages is about
20,000 to 25,000 messages per day and we were able to support this under Solaris
8/pmdf 6 with few problems.

When we initially upgraded to Solaris 8 and pmdf 6.0 in January 2001, we ran into
the SleepyCat problems almost daily. After reverting to Pmdf 5.2, we replaced our
software managed raid array (Volume manager 3.0.4 and Vx/fs 3.3.3) with a hardware
raid card. The hardware raid card freed up cpu cycles significantly and probably
some memory which is a little tight on our system. Shortly after this, we
upgraded pmdf back to 6.0 and have had few problems. Off the top of my head, I'd
say things have hung up three or four times since the middle of March, often after
I had changed our configuration and had done a pmdf cnbuild/pmdf restart.

Prior to replacing the software raid array, our system was cpu bound, io bound and
tight on memory (we usually have lots of IMAP connections to the server).
Switching to hardware raid significantly reduced cpu load and reduced io
bottlenecks; since then Pmdf 6 basically runs well on our system.

I wrote a message to info-pmdf about this on April 23rd (Re: 6.0 and Solaris
Confusion.. Should I upgrade?) which contains a few more details about this -
should be in the archives:
(http://www.pmdf.process.com/scripts/mxarchive/archive_search.com?TEXT=R416082-418676-mail%24archives%3A%5Binfo-pmdf%5Dinfo-pmdf.2001-04)

Scott Wood
Drew University

Richard Loken wrote:

--
Scott Wood, System Administrator swo...@drew.edu
Drew University, Madison NJ 07940 (973)-408-3657


ned+in...@mauve.mrochek.com

unread,
Jun 20, 2001, 10:43:52 AM6/20/01
to
> > I realise you guys want to go with a 'standard' database system, and
> > SleepyCat performs very well .... normally. But, wouldn't it be better
> > for your reputation and product stability to go either back to the
> > databases used in 5.2 or another standard like ndbm or gdbm or
> > something??

I'll comment further on the idea of using the V5.2 code below. As for ndbm or
gdbm, they do not support the necesssary features for the use PMDF makes of a
database, so they simply cannot be used. There are various relational
databases, but relational stuff would be serious overkill for what PMDF wants
to do.

> > I would really love a stable PMDF, I really would.

> So would we, Dave. I should probably avoid rambling replies like this
> one in a public forum, but it was asked publicly, so....

> From what they've told me, Innosoft found that the code they were


> using did not scale well for sites with higher throughput.

It also had its own share of other problems. One of the reasons for changing to
Sleepycat was the fact that we had any number of customers were were quite
unhappy with our own database code.

In effect switching to Sleepycat just changed one group of problems for
another. Switching back would just reverse the tradeoff. Worse, some of the
problems with the old database code are likely to be structural in nature and
hence nearly impossible to solve.

> So they
> decided to replace the home-grown code in favor of something better.
> They chose SleepyCat.

Actually, it more or less choose us. If you restrict yourself to directly
callable, simple (not relational), mulit-platform, thread-safe databases that
support record locking, you pretty much end up with Sleepycat.

> Given that I wasn't involved with PMDF at the
> time, I don't know exactly why they did choose SleepyCat, but I'm sure
> they wouldn't have gone that route if they'd believed ndbm or gdbm or
> something else had been better. (In fact, from the web searches we've
> done on the topic, the consensus in the UNIX world seems to be that
> SleepyCat is the best, even though it's far from perfect.)

Also true.

> So, rather than force the UNIX customers to suffer through the
> problems for another year or so while we drop the SleepyCat stuff and
> try something else, we're trying to concentrate on stabilizing what's
> there as quickly as we can. It's a frustrating situation for all of
> us (the problems existed, and went unaddressed, long before PMDF was
> transferred to Process Software), but we're working quite diligently
> on stabilizing the inherited problems as quickly as we can.

Other possibilities include using a server process that does all the database
stuff and which other processes call. Or there could be a single, separate
database thread that gets started in each PMDF process.

However, it is worth noting that we've basically given on databases entirely in
the iMS MTA. They are simply too unreliable, too inflexible, and too slow.

The most problematic database in the MTA is the queue cache. This is because it
is read and written by multiple threads in multiple processes. Other databases
like the alias database are typically read only and hence are much less of a
problem. In iMS the queue cache has been completely replaced by an in-memory
data structure maintained by the job controller. In additon to having
eliminated all the queue cache problems, the resulting increase in throughput
is nothing short of remarkable. And things like exponential backoff on message
retry become possible.

Ironically, the new job controller code was constructed during the PMDF V6.0
development process. But we decided it was too risky to go with such an
untested approach, and that a known quantity like Sleepycat was safer.

So much for the conservative approach.

Ned

David Powicki

unread,
Jun 20, 2001, 11:25:34 AM6/20/01
to

Scott,

Thanks for your message. I had seen you previous posting and was
somewhat encouraged that your earlier problems appeared to be related to
your software RAID configuration. Now, to hear that you continue to
have problems, albeit minor ones, with Solaris 8 and pmdf 6.0 reinforces
the idea that we will not be going to PMDF 6 until the Sleepycat
problems are resolved.

We run two mail hubs (E250s each with dual 450 CPUs and 1 GB ram) with
PMDF 5.2 and Solaris 7 (yes, Solaris 7). Each machine does 70-90K
messages per day and we have only had a problems with hung machines once
or twice in the last 2 years. Knock on wood, one of them has been up
for 497 days with PMDF having been restarted only a few times to change
configs. From what we are hearing 6.0 is ~safe for low traffic sites,
but certainly not for sites with the type of message volume we run...

David

--

Hunter Goatley

unread,
Jun 20, 2001, 1:50:14 PM6/20/01
to
> This historical perspective prompts me to raise the question of why not
> allow an option to disable the use of the queue cache and just use the
> file system again. On VMS this would be death on a system under load,
> but on UNIX not nearly as bad. And if the system was stable it might
> be an acceptable alternative. Not sure how much of PMDF would be
> affected by such an option, but I would put it forward for your
> consideration.

Thanks for the suggestion. We're also following up on Ned's comments.

Bill MacAllister

unread,
Jun 20, 2001, 1:49:59 PM6/20/01
to

--On Tuesday, June 19, 2001 6:10 PM -0500 Hunter Goatley
<goath...@goatley.com> wrote:

> rom what they've told me, Innosoft found that the code they were

> using did not scale well for sites with higher throughput. So they


> decided to replace the home-grown code in favor of something better.

> They chose SleepyCat. Given that I wasn't involved with PMDF at the


> time, I don't know exactly why they did choose SleepyCat, but I'm sure
> they wouldn't have gone that route if they'd believed ndbm or gdbm or
> something else had been better. (In fact, from the web searches we've
> done on the topic, the consensus in the UNIX world seems to be that
> SleepyCat is the best, even though it's far from perfect.)

One of the factors in deciding to use SleepyCat was that the Critical
Angle LDAP server, that later became IDDS, used it. So, there was a
body of understanding of SleepyCat at Innosoft. Of course, what was
not anticipated was stress that the queue cache would put on SleepyCat.
Remember, the queue cache was introduced to releave the performance
problems associated with using the RMS directory structure as in index
into messages in the queue. Thus, the volatility of the queue cache
was probably something that the designers of SleepyCat never expected
to see, in my opinion, much to their discredit. A database that loses
data or locks up under load is hardly useful. My motto is "works
first, fast later".

This historical perspective prompts me to raise the question of why not
allow an option to disable the use of the queue cache and just use the
file system again. On VMS this would be death on a system under load,
but on UNIX not nearly as bad. And if the system was stable it might
be an acceptable alternative. Not sure how much of PMDF would be
affected by such an option, but I would put it forward for your
consideration.

Bill

+-------------------------------------------------------
| Bill MacAllister, Senior Programmer
| PRIDE Industries
| 10030 Foothills Blvd., Dept 1150
| Roseville, California 95747
| Phone: +1 916-788-2402

Bill MacAllister

unread,
Jun 20, 2001, 2:08:25 PM6/20/01
to

--On Wednesday, June 20, 2001 12:50 PM -0500 Hunter Goatley
<goath...@goatley.com> wrote:

>> This historical perspective prompts me to raise the question of why
>> not allow an option to disable the use of the queue cache and just
>> use the file system again. On VMS this would be death on a system
>> under load, but on UNIX not nearly as bad. And if the system was
>> stable it might be an acceptable alternative. Not sure how much of
>> PMDF would be affected by such an option, but I would put it forward
>> for your consideration.
>

> Thanks for the suggestion. We're also following up on Ned's comments.

^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
Generally works for me too.

An in memory structure sounds wonderful. This is what some high volume
sites on VMS have done for years with ram disks, i.e. move the queue
cache database to the ram disk and rebuild it at every boot. An in
memory structure is a bit problematic for VMS clusters. Of course, iMS
doesn't give a hoot about VMS.

ned+in...@mauve.mrochek.com

unread,
Jun 20, 2001, 3:22:22 PM6/20/01
to
> An in memory structure sounds wonderful. This is what some high volume
> sites on VMS have done for years with ram disks, i.e. move the queue
> cache database to the ram disk and rebuild it at every boot. An in
> memory structure is a bit problematic for VMS clusters. Of course, iMS
> doesn't give a hoot about VMS.

True, but clusters in general we do care about. And Sun's newest clustering
software is very similar to VMS in the functionality it provides. In
particular, the necessary file system semantics appear to be present. (About
time.)

As a result we've thought about this, and the way to do it appears to be
something akin to the way the VMS queue manager works. (Surprise, surprise.)
That is, the job controller runs on only one cluster node with failover to
other nodes and processes on other machines talk to it over some sort of
network link.

Ned

ned....@mrochek.com

unread,
Jun 20, 2001, 6:51:22 PM6/20/01
to
> It would be interesting to know if Esys (Messaging Direct?) and Cyrus Imap
> are having database problems with SleepyCat.

I don't know anything about Messaging Direct's use. However, my understanding
of the Cyrus Imap use of Sleepycat is that it is:

(a) A fairly recent addition to the server. As such, information about it
may not be based on extensive experience.
(b) Read intensive, that is, information is read frequently but modified or
deleted rarely. In contrast, the usage in PMDF is basically 1:1:1
write/read/delete.
(c) Done from multiple single threaded processes. PMDF uses multiple
multi-threaded processes. And given the way Sleepycat works, this could
easily result in different database organizations, different use of
shared memory, and completely disjoint code paths.

I should also add that while we're moving away from Sleepycat databases in the
iMS MTA, the iMS store still uses them. I would characterize its style of use
as closer to PMDF than Cyrus IMAP. And although this is a different code base
developed by a different group of engineers, I hear that it has proved to be an
ongoing source of problems.

The bottom line is I doubt if the Cyrus experience is as good a source of
insight into PMDF's use of Sleepycat as you might think. Additionally, I'd say
that people who use Sleepycat mostly for reads or in an environment that
involves multiple processes or multiple threads but not both have far fewer
problems.

> I also speculate that these problems are more likely to show up with increased
> volumes of traffic. After I fixed my kernel quota parameters I have not seen
> any SCDB errors and the worst time for such errors was when the system was
> hung over the Easter weekend and we had a backlog of thousands of messages,
> I spent the morning rebuilding and recorrupting the cache until the backlog
> was gone and then the system hehaved quite well again.

The iMS store experience tends to confirm that volume is a factor, although
probably not as much of a factor as how the database is used.

Ned

0 new messages