Symptoms -
In large Microsoft Systems Management Server (SMS) 2003 hierarchies that have many sites, site-to-site replication slows down.
The volume of files may be larger than you expect in the following folders on a site server:
Sms\Inboxes\Schedule.box
Sms\Inboxes\Schedule.box\tosend
Sms\Inboxes\Replmgr.box\ready
These files represent site-to-site replication data that has been queued for processing by several components of the SMS_EXECUTIVE service. A baseline for the site is required to determine whether the counts are larger than expected. Large queues of replication information are occasionally expected. These large queues are typical when specific conditions exist.
Note The baseline is defined here as some historical measure of the volume of files in the Inboxes folder structure.
The following conditions can cause backlog scenarios:
Network or other infrastructure issues prevent the sender component from completing pending replication work.
Poor disk performance or slow I/O occurs because of a contention for disk resources.
SMS bandwidth restrictions limit the throughput of the sender component. This behavior keeps more send requests and jobs around for longer periods.
When addresses are unavailable, the SMS Scheduler component cannot schedule send requests by using the sender for the given address. This issue delays the part of the work that is associated with scheduling the send request until the address is available.
Distributing many or large packages in a short time creates a high load on the components that are involved in site-to-site replication.
Overly aggressive schedules exist for discovery data generation, inventory collection, collection evaluation, and so on.
In a hierarchy that has three or more tiers, middle-tier sites that have many child sites handle larger volumes of jobs and replication objects. This behavior occurs because of site-to-site replication routing. The load of a middle-tier site is increased for each child site that is attached. Therefore, reducing the number of attached sites can, in some cases, reduce this load.
Sites are removed from the hierarchy incorrectly.
In most cases, when the conditions that cause significant replication queuing have been corrected or when these conditions have subsided, the queued replication data is processed and then cleared.
Cause:
When the SMS Scheduler component is processing large quantities of active jobs and send requests, the throughput of the Scheduler component begins to slow. This behavior occurs because of a corresponding increase in processing overhead for the increased quantities of objects.
In some instances, if a large enough queue of data is formed, it can take days or even weeks to be completely processed. The time that is required to process the queued data depends on the many variables that affect replication performance in the hierarchy and in the environment. These variables include disk I/O performance, network speeds, bandwidth restrictions, size of queued data, and object count. When a large queue of backlogged replication data has been formed, adding additional loads increases the time that is required for all data to be processed.
In most cases, the appropriate action for a large backlog of replication data is to first correct any issue that may be preventing processing of replication data. Next, you may have to reduce the quantity of site-to-site replication traffic. Finally, make sure that the SMS_EXECUTIVE service can run uninterrupted to complete processing in a timely manner. Service restarts can add significant overhead. Limiting SMS_EXECUTIVE service restarts is important because the initialization work for the SMS Scheduler component is proportional to the number of jobs, send requests, and routing requests that are currently queued for processing.
Note The SMS_EXECUTIVE service hosts the SMS Replication Manager, SMS Scheduler, and SMS Sender components.
June 16, 2009
How to change the credentials for the OpsMgr SDK Service and for the OpsMgr Config Service in Microsoft System Center Operations Manager 2007
Reference:
http://support.microsoft.com/kb/936220
Regards,
Atul
http://support.microsoft.com/kb/936220
Regards,
Atul
June 5, 2009
Antivirus software blocks script execution in System Center Operations Manager 2007
SYMPTOMS
In Microsoft System Center Operations Manager 2007, you may receive alerts that have a warning severity that resembles the following:
Script or Executable Failed to run
The process started at 10:41:22 AM failed to create System.PropertyBagData, no errors detected in the output. The process exited with 1
Command executed: "C:\WINDOWS\system32\cscript.exe" //nologo "C:\Program Files\System Center Operations Manager 2007\Health Service State\Monitoring Host Temporary Files 73\3456\ScriptName.vbs"
Working Directory: C:\Program Files\System Center Operations Manager 2007\Health Service State\Monitoring Host Temporary Files 73\3456\
One or more workflows were affected by this.
CAUSE
This problem occurs because some antivirus software blocks Visual Basic scripts or Java scripts.
RESOLUTION
To resolve this problem, verify that your antivirus software is not blocking scripts from running.
In Microsoft System Center Operations Manager 2007, you may receive alerts that have a warning severity that resembles the following:
Script or Executable Failed to run
The process started at 10:41:22 AM failed to create System.PropertyBagData, no errors detected in the output. The process exited with 1
Command executed: "C:\WINDOWS\system32\cscript.exe" //nologo "C:\Program Files\System Center Operations Manager 2007\Health Service State\Monitoring Host Temporary Files 73\3456\ScriptName.vbs"
Working Directory: C:\Program Files\System Center Operations Manager 2007\Health Service State\Monitoring Host Temporary Files 73\3456\
One or more workflows were affected by this.
CAUSE
This problem occurs because some antivirus software blocks Visual Basic scripts or Java scripts.
RESOLUTION
To resolve this problem, verify that your antivirus software is not blocking scripts from running.
OpsMgr 2007: Files and Folders starting with "Program" causing unmonitored Agent
Symptom
The Operations Manager Service Pack 1 (SP1)Â Agent or Management Server may be shown as greyed out in the Operations Manager console and the following events may be logged in the event log:
Event ID: 10000
Source: DCOM
Description: Unable to start a Dcom Server: {}. The error: description> Happened while starting this command: -Embedding
regarding monitoring host
-and-
Event Type: Error
Event Source: HealthService
Event Category: Health Service
Event ID: 1102
Description: Rule/Monitor
"Microsoft.SystemCenter.DiscoveryHealthServiceCommunication" running for instance
"" with id:"{38696FAA-2A83-6068-B008-DB43D49FB879}" cannot be
initialized and will not be loaded. Management group ""
Cause
Computers having files or folder that start with "Program" on the root drive may not be monitored. All workflows fail when file "c:\Program" is present on the machine. This happens because HealthService.exe is unable to start MonitoringHost.exe.
Workaround Information
To resolve this issue, delete or rename the file or folder named Program on the affected computer.
The Operations Manager Service Pack 1 (SP1)Â Agent or Management Server may be shown as greyed out in the Operations Manager console and the following events may be logged in the event log:
Event ID: 10000
Source: DCOM
Description: Unable to start a Dcom Server: {
regarding monitoring host
-and-
Event Type: Error
Event Source: HealthService
Event Category: Health Service
Event ID: 1102
Description: Rule/Monitor
"Microsoft.SystemCenter.DiscoveryHealthServiceCommunication" running for instance
"
initialized and will not be loaded. Management group "
Cause
Computers having files or folder that start with "Program" on the root drive may not be monitored. All workflows fail when file "c:\Program" is present on the machine. This happens because HealthService.exe is unable to start MonitoringHost.exe.
Workaround Information
To resolve this issue, delete or rename the file or folder named Program on the affected computer.
The W3WP.exe process crashes when the Anonymous authentication is disabled on the IISADMPWD virtual directory
SYMPTOMS
When a user's password is expired, you can use the Anonymous user account to change the expired password through the achg.asp file even when the Anonymous authentication is disabled on the IISADMPWD virtual directory.In this situation, if the AnonymousUserName and the AnonymousUserPass metabese properties are inconsistent or the "denied access this computer from the network" policy is applied for the Anonymous user, the Anonymous user cannot log on the server and an access violation occurs. In addition, the W3WP.exe process crashes.
RESOLUTION
To avoid this effect, use one of the following methods:
Set correct AnonymousUserName and AnonymousUserPass metabese properties or disable the "denied access this computer from the network" policy for anonymous user.
Separate the Application pool for the IISADMPWD virtual directory. Note A user may receive the 403.18 error when the request is redirected to the IISADMPWD password change pages, and the password cannot be changed through IISADMPWD. However, the W3WP.exe process does not crash.Note These settings violate the Internet Information Services (IIS) requirements that are described in the following Microsoft Knowledge Base:
812614Â (http://kbalertz.com/Feedback.aspx?kbNumber=812614/ ) Default permissions and user rights for IIS 6.0
Steps to reproduce this problemTo reproduce the problem, follow these steps:
On the IISADMPWD virtual directory, set an incorrect password in the AnonymousUserPass metabase for the IUSR account, or apply the "denied access this computer from the network" policy for the IUSR account.
Create a new local or domain user and enable "change their password the next time that the user logs on." This means the user's password is expired.
Disable Anonymous authentication for IISADMPWD.
Enable Basic authentication or Integrated Windows authentication for IISADMPWD.
Create a new TEST virtual directory that is enabled Basic or Integrated Windows authentication. When you access the TEST virtual directory, you will be redirected to the aexp3.asp Web page because the password is expired. If you enter an old password and a new password, and then click OK, the dialog box for Basic authentication appears. If you enter the old password, you will experience the symptoms that are described in the "Symptoms" section.
When a user's password is expired, you can use the Anonymous user account to change the expired password through the achg.asp file even when the Anonymous authentication is disabled on the IISADMPWD virtual directory.In this situation, if the AnonymousUserName and the AnonymousUserPass metabese properties are inconsistent or the "denied access this computer from the network" policy is applied for the Anonymous user, the Anonymous user cannot log on the server and an access violation occurs. In addition, the W3WP.exe process crashes.
RESOLUTION
To avoid this effect, use one of the following methods:
Set correct AnonymousUserName and AnonymousUserPass metabese properties or disable the "denied access this computer from the network" policy for anonymous user.
Separate the Application pool for the IISADMPWD virtual directory. Note A user may receive the 403.18 error when the request is redirected to the IISADMPWD password change pages, and the password cannot be changed through IISADMPWD. However, the W3WP.exe process does not crash.Note These settings violate the Internet Information Services (IIS) requirements that are described in the following Microsoft Knowledge Base:
812614Â (http://kbalertz.com/Feedback.aspx?kbNumber=812614/ ) Default permissions and user rights for IIS 6.0
Steps to reproduce this problemTo reproduce the problem, follow these steps:
On the IISADMPWD virtual directory, set an incorrect password in the AnonymousUserPass metabase for the IUSR account, or apply the "denied access this computer from the network" policy for the IUSR account.
Create a new local or domain user and enable "change their password the next time that the user logs on." This means the user's password is expired.
Disable Anonymous authentication for IISADMPWD.
Enable Basic authentication or Integrated Windows authentication for IISADMPWD.
Create a new TEST virtual directory that is enabled Basic or Integrated Windows authentication. When you access the TEST virtual directory, you will be redirected to the aexp3.asp Web page because the password is expired. If you enter an old password and a new password, and then click OK, the dialog box for Basic authentication appears. If you enter the old password, you will experience the symptoms that are described in the "Symptoms" section.
June 1, 2009
Operations Manager 2007 Design Tips
The following are some tips to consider when designing your Operations Manager 2007 infrastructure.
1. Always setup a minimum of 1 RMS and 1 MS. Do not have agents report directly to the RMS. remember that the RMS functions to distribute configuration information to all MS. Having additional load on to this process is not recommended. Besides, with this, you'll have a failover scenario in place.
2. 3-node clusters for RMS is not supported
3. To have a affordable failover strategy for your Operations DB, use SQL Log shipping. Unfortunately, DB Mirroring is an unsupported method.
4. When dealing with multi-site monitoring (branches), use a Gateway Server instead of a MS. Have MS in close proximity with your SQL Server. Why? Cause whenever MS needs to write data, it establishes a SQL ODBC connectivity. This takes up resources and the data is uncompressed. By using a GWS, data is compressed and the connection to a MS is always connected.
5. Have a dedicated MS for reporting from a GWS. Do not have other agents reporting to the same MS as a GWS. Reason is that Management Servers divide their processes by number of connections. Let's say that you have 10 servers reporting to the GWS. When the MS receives that connection, it is treated as 1. If you had an additional of 10 servers reporting to that MS, the MS will divide its performance 11 ways. You would then see a significant performance drop for the servers handled by the GWS. If GWS is the only one connected to the MS, it will be given the full 100%.
6. The RMS consumes CPU and RAM as its core process. So bulk up on these
7. Use 64-bit for the RMS so that there are opportunities to scale beyond 4GB of RAM
8. There is a Datawarehouse Grooming tool found in the Resource Kit that will help trim down the size of the Operations DW
9. Support for SQL 2008 will be around the August 2008 timeframe or SP2. This will be cool cause there will be no dependency on IIS
10. Each GWS can support up to 800 Agents with the SP1
1. Always setup a minimum of 1 RMS and 1 MS. Do not have agents report directly to the RMS. remember that the RMS functions to distribute configuration information to all MS. Having additional load on to this process is not recommended. Besides, with this, you'll have a failover scenario in place.
2. 3-node clusters for RMS is not supported
3. To have a affordable failover strategy for your Operations DB, use SQL Log shipping. Unfortunately, DB Mirroring is an unsupported method.
4. When dealing with multi-site monitoring (branches), use a Gateway Server instead of a MS. Have MS in close proximity with your SQL Server. Why? Cause whenever MS needs to write data, it establishes a SQL ODBC connectivity. This takes up resources and the data is uncompressed. By using a GWS, data is compressed and the connection to a MS is always connected.
5. Have a dedicated MS for reporting from a GWS. Do not have other agents reporting to the same MS as a GWS. Reason is that Management Servers divide their processes by number of connections. Let's say that you have 10 servers reporting to the GWS. When the MS receives that connection, it is treated as 1. If you had an additional of 10 servers reporting to that MS, the MS will divide its performance 11 ways. You would then see a significant performance drop for the servers handled by the GWS. If GWS is the only one connected to the MS, it will be given the full 100%.
6. The RMS consumes CPU and RAM as its core process. So bulk up on these
7. Use 64-bit for the RMS so that there are opportunities to scale beyond 4GB of RAM
8. There is a Datawarehouse Grooming tool found in the Resource Kit that will help trim down the size of the Operations DW
9. Support for SQL 2008 will be around the August 2008 timeframe or SP2. This will be cool cause there will be no dependency on IIS
10. Each GWS can support up to 800 Agents with the SP1
Subscribe to:
Posts (Atom)